Multi-mode visual communication display device

By combining the multimodal sensing module, the core control module, and the power management module, the problems of inaccurate data fusion, single display mode, high energy consumption, and poor user interaction experience in multimodal visual communication display devices are solved, achieving a multimodal display effect with accurate adaptation, high stability, low energy consumption, and natural interaction.

CN121879703APending Publication Date: 2026-04-17吴佳忆
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
吴佳忆
Filing Date
2026-01-09
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing multimodal visual communication display devices have simple and crude data fusion methods, fail to consider environmental complexity and user interaction activity, have limited display mode switching, lack modal conflict resolution mechanisms, have limited energy consumption control, and have limited user interaction feedback forms, resulting in poor adaptability, poor experience, weak stability, and high power consumption.

Method used

The system employs a multimodal perception module to collect environmental and user interaction data. Through the core control module, a multimodal data fusion unit and a display strategy optimization unit are built in, and weights are dynamically allocated in conjunction with an attention mechanism. A modal conflict mediation unit and a power management module are set up to achieve dynamic energy consumption adjustment. A haptic feedback component is also added to support a wireless communication module.

Benefits of technology

It achieves accurate fusion of multimodal data, adaptive optimization of display strategies, rapid resolution of modal conflicts, dynamic adjustment of energy consumption, improves the naturalness of user interaction experience and the stability and battery life of the device, and enhances the versatility and scalability of the device.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121879703A_ABST
    Figure CN121879703A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of visual communication, and discloses a multi-modal visual communication display device, which comprises a multi-modal sensing module used for collecting environment parameter data and user interaction data, and the environment parameter data at least comprises environment light intensity and environment color temperature; the user interaction data at least comprises a user gesture action, a user sight direction and a voice instruction; the wireless communication module is integrated, various communication protocols such as Wi-Fi, Bluetooth and NFC are supported, remote control and data interaction with an external terminal can be achieved, meanwhile, different application scenes such as an intelligent cabin, intelligent office display and public information display can be flexibly adapted through the modular design and the self-adaptive display strategy, and the application range is wide. Compared with the defect of single scene adaptability in the prior art, the method has the advantages that the universality and expandability of the device are remarkably improved, and the customization development cost in different scenes is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of visual communication technology, and more particularly to a multimodal visual communication display device. Background Technology

[0002] With the rapid development of display technology and human-computer interaction technology, multimodal visual communication display devices have been widely used in various fields such as smart cockpits, smart offices, and public information displays. Their core requirement is to achieve more accurate and efficient visual information transmission and more natural user interaction. Existing multimodal visual communication display devices typically integrate multiple display components and sensing components, attempting to achieve optimized display effects through multimodal data collaboration.

[0003] However, existing technologies still have many shortcomings that urgently need to be addressed: First, the multimodal data fusion method is simple and crude, mostly using a weighted summation method with fixed weights, without considering the dynamic changes in environmental complexity and user interaction activity. This leads to a mismatch between the fusion decision information and the actual use scenario, which in turn affects the adaptability of the display strategy. Secondly, the display mode switching often relies on a single environmental parameter (such as only based on ambient light intensity) without taking into account the user's interaction state (such as eye focus, gestures), which fails to meet the user's personalized visual experience needs and is prone to conflict between display mode and user needs. Third, the lack of an effective modal conflict resolution mechanism means that when the display decisions corresponding to different modal data contradict each other, the conflict cannot be resolved quickly and accurately, resulting in poor device operation stability. Fourth, the energy consumption control strategy is simplistic and fails to dynamically adjust energy consumption based on user interaction status and module operation requirements. High power consumption still exists in standby or low-load scenarios, affecting the device's battery life. Fifth, the user interaction feedback is monotonous, mostly relying on visual feedback and lacking multi-dimensional feedback such as tactile feedback, resulting in an inconsistent and unnatural user interaction experience.

[0004] Based on the shortcomings of the existing technologies, there is an urgent need for a multimodal visual communication display device that can achieve dynamic and accurate fusion of multimodal data, adaptive optimization of display strategies, effective resolution of modal conflicts, dynamic adjustment of energy consumption, and improvement of interactive experience, so as to solve the problems of poor adaptability, poor experience, weak stability, and high power consumption in the existing technologies. Summary of the Invention

[0005] In order to overcome the shortcomings of the prior art, one of the objectives of the present invention is to provide a multimodal visual communication display device.

[0006] One of the objectives of this invention is achieved through the following technical solution: A multimodal visual communication display device, comprising: A multimodal perception module is used to collect environmental parameter data and user interaction data. The environmental parameter data includes at least ambient light intensity and ambient color temperature, and the user interaction data includes at least user gestures, user gaze direction, and voice commands. The core control module is communicatively connected to the multimodal perception module. The core control module has a built-in multimodal data fusion unit and a display strategy optimization unit. The multimodal data fusion unit is used to extract features and perform weighted fusion on the collected environmental parameter data and user interaction data to obtain fused decision information. The display strategy optimization unit generates target display control instructions based on the fused decision information. A multimodal display execution module is communicatively connected to the core control module. The multimodal display execution module includes a main display panel, an auxiliary display component, and a light adjustment component. The main display panel is used to present core visual content, the auxiliary display component is used to supplement and enhance visual information, and the light adjustment component is used to adjust the display light intensity and color temperature according to the target display control command. The power management module is electrically connected to the multimodal sensing module, the core control module, and the multimodal display execution module, respectively, and is used to provide stable power supply to each module and realize dynamic energy consumption adjustment.

[0007] Furthermore, the multimodal sensing module includes: An environmental sensing unit, comprising a light intensity sensor and a color temperature sensor, for real-time acquisition of ambient light intensity data and ambient color temperature data, respectively; The user interaction sensing unit includes an image acquisition unit, a voice acquisition unit, and a gaze tracking sensor. The image acquisition unit is used to capture images of user gestures, the voice acquisition unit is used to receive user voice commands, and the gaze tracking sensor is used to detect the user's gaze focus position.

[0008] Furthermore, the fusion process of the multimodal data fusion unit is as follows: First, the data of each modality is preprocessed, noisy data is filtered and standardized; then, an attention mechanism is used to assign dynamic weights to different modalities, with the weight coefficient of environmental parameter data ranging from 0.3 to 0.5 and the weight coefficient of user interaction data ranging from 0.5 to 0.7; finally, the fusion decision information is obtained by weighted summation.

[0009] Furthermore, the display strategy optimization unit incorporates multiple display modes, including high-definition color mode, low-power eye protection mode, and enhanced contrast mode. The display strategy optimization unit automatically switches the display mode based on the ambient light intensity threshold and the user's gaze state in the fusion decision information: when the ambient light intensity is ≥500 lux, it switches to enhanced contrast mode; when the ambient light intensity is <100 lux, it switches to low-power eye protection mode; and when the user's gaze focus time exceeds 30 seconds, it maintains high-definition color mode.

[0010] Furthermore, the main display panel adopts an AMOLED flexible display panel, and the auxiliary display component includes several distributed micro-display units. The micro-display units are evenly distributed in the edge area of ​​the main display panel and are used to display prompt information, navigation information or supplementary data. The light adjustment component includes an LED backlight module and a color temperature adjustment filter. The LED backlight module adopts zone control technology, and the brightness of each zone can be adjusted independently.

[0011] Furthermore, the core control module also has a built-in modal conflict resolution unit. When different modal data collected by the multimodal perception module cause decision conflicts, the modal conflict resolution unit resolves the conflicts based on preset priority rules, in which user gesture commands have higher priority than voice commands, and user interaction data has higher priority than environmental parameter data.

[0012] Furthermore, the power management module includes a power conversion unit, an energy consumption monitoring unit, and an intelligent switch unit. The energy consumption monitoring unit collects power consumption data of each module in real time. When the device is in standby mode and there is no user interaction for more than 5 minutes, the intelligent switch unit automatically turns off the power supply to the auxiliary display component and some environmental sensing units, and only retains the low-power operation of the core sensing component.

[0013] Furthermore, the image acquisition device uses a binocular camera to recognize the user's three-dimensional gestures through a stereo vision algorithm. The recognizable gestures include clicking, swiping, zooming, and rotating, with a recognition accuracy of ≥95%. The gaze tracking sensor collects the coordinates of the user's eye feature points and calculates the focus coordinates of the user's gaze on the main display panel, with a positioning error of ≤5mm.

[0014] Furthermore, it also includes a wireless communication module, which is connected to the core control module and supports Wi-Fi, Bluetooth and NFC communication protocols, enabling multimodal data interaction with external terminals and remote control command reception.

[0015] Furthermore, the multimodal display execution module also includes a haptic feedback component, which is connected to the core control module. When the user triggers an interactive operation through gestures, the haptic feedback component outputs vibration feedback of different intensities according to the type of interaction. The vibration intensity is adjustable in 3-5 levels.

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention, by embedding a multimodal data fusion unit in the core control module, dynamically allocates the weights of environmental parameter data and user interaction data using an attention mechanism. It can adjust the weight coefficients in real time according to the complexity of the environment (such as strong light or weak light environment) and the activity of user interaction (such as active gestures, voice interaction or no interaction), thus solving the problem of mismatch between decision information and scene caused by fixed weight fusion in the prior art. This makes the fused decision information more in line with the actual use scenario, thereby ensuring the accurate adaptation of the display strategy.

[0017] 2. The display strategy optimization unit of the present invention combines environmental parameter data (light intensity, color temperature) and user interaction data (eye focus, gesture action) to realize multi-dimensional display mode switching, covering a variety of scenario-based modes such as high-definition color mode, low-power eye protection mode, and enhanced contrast mode, which can accurately match the visual needs of different environments and user states.

[0018] 3. By setting up a modal conflict mediation unit, this invention establishes a priority-based conflict resolution rule (user interaction data has higher priority than environmental parameter data, and gesture commands have higher priority than voice commands). This can quickly resolve display decision conflicts corresponding to different modal data. Compared with the device operation disorder caused by the lack of a conflict mediation mechanism in the prior art, this invention effectively improves the stability and reliability of multimodal collaborative operation and ensures the accurate execution of display control commands.

[0019] 4. The power management module of this invention integrates an energy consumption monitoring unit and an intelligent switching unit, enabling it to collect power consumption data of each module in real time and dynamically adjust the power supply strategy based on user interaction status. When in standby mode and without user interaction for more than a preset time, it automatically shuts off power to non-core modules, retaining only the core sensing components for low-power operation. This solves the high power consumption problem caused by the single energy consumption control in existing technologies. Actual testing shows that the power consumption of this device in standby low-power mode can be reduced to below 50mW, extending the battery life by more than 30% compared to existing similar devices, making it particularly suitable for portable or battery-powered applications.

[0020] 5. This invention adds a tactile feedback component to the multimodal display execution module, outputting vibration feedback of varying intensity according to different gesture interaction types. This achieves synergy between visual and tactile feedback. Compared to existing technologies that rely solely on visual feedback, this invention enables users to intuitively perceive the effectiveness of interactive operations through touch, improving the naturalness and coherence of human-computer interaction and reducing the learning cost of interaction.

[0021] 6. This invention integrates a wireless communication module that supports multiple communication protocols such as Wi-Fi, Bluetooth, and NFC, enabling remote control and data interaction with external terminals. Furthermore, through modular design and adaptive display strategies, it can flexibly adapt to different application scenarios such as smart cockpits, smart office displays, and public information displays. Compared to the limitations of existing technologies with their limited adaptability, this invention significantly improves the versatility and scalability of the device, reducing the cost of customized development for different scenarios.

[0022] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described in detail below with reference to the accompanying drawings. Attached Figure Description

[0023] Figure 1 This is a diagram illustrating the overall structure and workflow of this embodiment; Figure 2 This is a flowchart of the data fusion and display control process in this embodiment. Detailed Implementation

[0024] The present invention will now be further described in conjunction with the accompanying drawings and specific embodiments. It should be noted that, without conflict, the various embodiments or technical features described below can be arbitrarily combined to form new embodiments.

[0025] It should be noted that when a component is described as "fixed to" another component, it can be directly on the other component or may have a component in between. When a component is considered "connected to" another component, it can be directly connected to the other component or may have a component in between. When a component is considered "set on" another component, it can be directly set on the other component or may have a component in between. The terms "vertical," "horizontal," "left," "right," and similar expressions used in this document are for illustrative purposes only.

[0026] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0027] I. The multimodal visual communication display device in this embodiment includes a multimodal sensing module, a core control module, a multimodal display execution module, a power management module, and a wireless communication module. These modules are integrated and connected via a PCB board. Specific hardware selections are as follows: 1. Multimodal sensing module: Environmental sensing unit: The BH1750 light intensity sensor (measurement range 0-65535 lux, accuracy ±20%) and TCS34725 color temperature sensor (measurement range 2000K-7000K) are selected. Both communicate with the core control module through the I2C bus. The user interaction sensing unit uses a binocular OV5640 camera as an image acquisition device (1080P resolution, 30fps frame rate), an SGM3770 microphone array as a voice acquisition device (supporting 8kHz-48kHz sampling rate), and a PupilLabsCore eye-tracking sensor (positioning error ≤3mm). The camera and eye-tracking sensor are connected to the core control module via a MIPI interface, and the microphone array is connected via an SPI interface.

[0028] 2. Core Control Module: The STM32H743VIT6 microcontroller is selected as the core chip (480MHz main frequency, supports multi-protocol communication). The built-in multimodal data fusion unit and display strategy optimization unit are implemented through embedded software programming; the modal conflict mediation unit is based on priority judgment logic written in C language and stored in the chip's Flash memory.

[0029] 3. Multimodal display execution module: Main display panel: Samsung A3 AMOLED flexible display panel (size 10.1 inches, resolution 2560×1600, contrast ratio 1,000,000:1). Auxiliary display components: Eight TDKCM1608 miniature OLED display units (0.96 inches in size, 128×64 resolution) are selected and evenly distributed around the edges of the main display panel, and connected to the core control module through the SPI interface; Light adjustment components: The backlight module adopts an ADSPro type local dimming LED backlight module (32 dimming zones) and an adjustable color temperature filter for liquid crystal (adjustment range 2700K-6500K). The backlight module is controlled by PWM signal and the filter is controlled by I2C bus. Haptic feedback component: The TIDRV2605 vibration motor driver chip is selected, paired with a 10mm miniature linear vibration motor, and connected to the core control module via I2C bus.

[0030] 4. Power Management Module: It adopts an RT9013 LDO regulator (output 3.3V / 2A) and a TPS61088 boost chip (input 3.7V-5V, output 5V / 3A). The energy consumption monitoring unit uses an INA219 current and voltage monitoring chip (measurement accuracy ±1%). The intelligent switching unit uses an AO3400 MOSFET. All components are connected to the functional modules through the power bus.

[0031] 5. Wireless communication module: The ESP32-WROOM-32E module is selected, which integrates Wi-Fi (802.11b / g / n), Bluetooth 5.0 and NFC functions. It communicates with the core control module through the UART interface, and the antenna is a built-in antenna on the PCB.

[0032] II. Each module is physically connected and transmits signals via a PCB board: The environmental sensing unit and user interaction sensing unit of the multimodal sensing module are connected to the corresponding interfaces of the core control module via I2C, MIPI, and SPI interfaces, respectively; the main display panel of the multimodal display execution module is connected to the core control module via the MIPI / SPI interface, and the auxiliary display component, light adjustment component, and haptic feedback component are connected to the core control module via SPI, PWM, and I2C interfaces, respectively; the output of the power management module is connected to the power input of each module via power copper foil, and the signal output of the energy consumption monitoring unit is connected to the ADC interface of the core control module; the wireless communication module is connected to the core control module via the UART interface to realize data interaction.

[0033] III. Device Working Process The operation process of the multimodal visual communication display device in this embodiment is as follows: 1. Startup and Initialization: After the device is powered on, the RT9013 regulator and TPS61088 boost chip of the power management module work to provide stable voltage for each module; the core control module STM32H743VIT6 chip starts up, executes the initialization program, configures the parameters of each interface, initializes the parameters of the multi-mode data fusion unit, display strategy optimization unit and mode conflict mediation unit, initializes the wireless communication module and automatically searches for available Wi-Fi networks or Bluetooth devices, and enters standby mode after initialization is completed.

[0034] 2. Multimodal Data Acquisition: The multimodal perception module begins real-time data acquisition: the BH1750 light intensity sensor and TCS34725 color temperature sensor of the environmental sensing unit acquire ambient light intensity and color temperature data every 100ms and transmit it to the core control module; the binocular OV5640 camera of the user interaction sensing unit captures a frame of user gesture image every 33ms (30fps), the SGM3770 microphone array acquires user voice commands in real time, and the PupilLabsCore gaze tracking sensor detects the user's gaze focus position every 50ms. The acquired user interaction data is transmitted to the core control module in real time.

[0035] 3. Multimodal Data Fusion: The multimodal data fusion unit of the core control module preprocesses the received environmental parameter data and user interaction data: it filters noise in the environmental data using a median filter algorithm, optimizes user gesture images and gaze data using a Gaussian filter algorithm, and then standardizes all data (mapping the data to the 0-1 range); subsequently, it uses an attention mechanism to allocate dynamic weights, adjusting the weight coefficients according to environmental complexity and user interaction activity. When the user actively interacts (such as gestures or voice commands), the weight coefficient for user interaction data is set to 0.6-0.7, and the weight coefficient for environmental parameter data is set to 0.3-0.4; when there is no active user interaction, the weight coefficient for both user interaction data and environmental parameter data is set to 0.5; finally, the fused decision information is obtained through a weighted summation formula (fused decision information = environmental parameter data × environmental weight + user interaction data × user weight).

[0036] 4. Display Strategy Optimization and Control: The display strategy optimization unit generates target display control instructions based on fused decision information. When the ambient light intensity in the fused decision information is ≥500 lux, an enhanced contrast mode control command is generated. The core control module controls the main display panel to increase the contrast (adjusted to 1200000:1) through the MIPIDSI interface, controls the LED backlight module of the light adjustment component to increase the brightness of each zone, and adjusts the color temperature to 6500K. When the ambient light intensity is less than 100 lux, a low-power eye protection mode control command is generated, and the main display panel reduces its brightness to 50 cd / m². 2 The LED backlight module turns off the backlight in some non-core areas, adjusts the color temperature to 2700K, and the auxiliary display components only retain the display of necessary prompts. When the user's gaze is focused for more than 30 seconds, a high-definition color mode control command is generated, the main display panel turns on 10-bit color depth display, the color saturation is adjusted to 100%, and the auxiliary display components simultaneously display supplementary data. When a user triggers an interactive operation (such as clicking or swiping) through gestures, the core control module controls the haptic feedback component to output vibration feedback of corresponding intensity according to the interaction type (click operation corresponds to level 1 vibration, swipe operation corresponds to level 2 vibration, zoom and rotate operation correspond to level 3 vibration).

[0037] 5. Modal conflict resolution: When different modal data collected by the multimodal perception module cause decision conflicts (such as a user's voice command requesting a switch to high-definition color mode, while the ambient light intensity <100 lux corresponds to a low-power eye protection mode), the modal conflict mediation unit activates a preset priority rule to resolve the conflict—since user interaction data has a higher priority than environmental parameter data, the high-definition color mode control command is ultimately executed; if there are multiple user interaction data conflicts (such as gesture commands and voice commands being inconsistent), the gesture command with higher priority will prevail.

[0038] 6. Dynamic Energy Consumption Adjustment and Data Interaction: The INA219 energy consumption monitoring unit of the power management module collects power consumption data of each module every 200ms and transmits it to the core control module. When the device is in standby mode and there is no user interaction for more than 5 minutes, the core control module controls the intelligent switch unit to turn off the power supply to the auxiliary display component, voice collector and some environmental sensing units, and only retains the low-power operation of the binocular camera, eye tracking sensor and core control module, reducing the power consumption to below 50mW. The wireless communication module can realize multimodal data interaction with external terminals (such as mobile phones and computers), upload the collected environmental parameter data and user interaction data to the external terminal, and receive remote control commands from the external terminal (such as remote switching of display mode) and transmit them to the core control module for execution.

[0039] The above embodiments are merely preferred embodiments of the present invention and should not be construed as limiting the scope of protection of the present invention. Any non-substantial changes and substitutions made by those skilled in the art based on the present invention shall fall within the scope of protection claimed by the present invention.

Claims

1. A multimodal visual communication display device, characterized in that, include: A multimodal perception module is used to collect environmental parameter data and user interaction data. The environmental parameter data includes at least ambient light intensity and ambient color temperature, and the user interaction data includes at least user gestures, user gaze direction, and voice commands. The core control module is communicatively connected to the multimodal perception module. The core control module has a built-in multimodal data fusion unit and a display strategy optimization unit. The multimodal data fusion unit is used to extract features and perform weighted fusion on the collected environmental parameter data and user interaction data to obtain fused decision information. The display strategy optimization unit generates target display control instructions based on the fused decision information. A multimodal display execution module is communicatively connected to the core control module. The multimodal display execution module includes a main display panel, an auxiliary display component, and a light adjustment component. The main display panel is used to present core visual content, the auxiliary display component is used to supplement and enhance visual information, and the light adjustment component is used to adjust the display light intensity and color temperature according to the target display control command. The power management module is electrically connected to the multimodal sensing module, the core control module, and the multimodal display execution module, respectively, and is used to provide stable power supply to each module and realize dynamic energy consumption adjustment.

2. The multimodal visual communication display device according to claim 1, characterized in that, The multimodal sensing module includes: An environmental sensing unit, comprising a light intensity sensor and a color temperature sensor, for real-time acquisition of ambient light intensity data and ambient color temperature data, respectively; The user interaction sensing unit includes an image acquisition unit, a voice acquisition unit, and a gaze tracking sensor. The image acquisition unit is used to capture images of user gestures, the voice acquisition unit is used to receive user voice commands, and the gaze tracking sensor is used to detect the user's gaze focus position.

3. The multimodal visual communication display device according to claim 1, characterized in that, The fusion process of the multimodal data fusion unit is as follows: First, the data of each modality is preprocessed, noisy data is filtered and standardized; then, an attention mechanism is used to assign dynamic weights to different modalities, with the weight coefficient of environmental parameter data ranging from 0.3 to 0.5 and the weight coefficient of user interaction data ranging from 0.5 to 0.7; finally, the fusion decision information is obtained by weighted summation.

4. The multimodal visual communication display device according to claim 1, characterized in that, The display strategy optimization unit has multiple built-in display modes, including high-definition color mode, low-power eye protection mode, and enhanced contrast mode. The display strategy optimization unit automatically switches the display mode according to the ambient light intensity threshold and the user's gaze state in the fusion decision information: when the ambient light intensity is ≥500 lux, it switches to enhanced contrast mode; when the ambient light intensity is <100 lux, it switches to low-power eye protection mode; and when the user's gaze focus time exceeds 30 seconds, it maintains high-definition color mode.

5. The multimodal visual communication display device according to claim 1, characterized in that, The main display panel uses an AMOLED flexible display panel, and the auxiliary display components include several distributed micro-display units. The micro-display units are evenly distributed in the edge area of ​​the main display panel and are used to display prompt information, navigation information or supplementary data. The light adjustment components include an LED backlight module and a color temperature adjustment filter. The LED backlight module adopts zone control technology, and the brightness of each zone can be adjusted independently.

6. The multimodal visual communication display device according to claim 1, characterized in that, The core control module also has a built-in modal conflict resolution unit. When different modal data collected by the multimodal perception module cause decision conflicts, the modal conflict resolution unit resolves the conflicts based on preset priority rules, in which user gesture commands have higher priority than voice commands, and user interaction data has higher priority than environmental parameter data.

7. The multimodal visual communication display device according to claim 1, characterized in that, The power management module includes a power conversion unit, an energy consumption monitoring unit, and an intelligent switch unit. The energy consumption monitoring unit collects power consumption data of each module in real time. When the device is in standby mode and there is no user interaction for more than 5 minutes, the intelligent switch unit automatically turns off the power supply to the auxiliary display component and some environmental sensing units, and only retains the low-power operation of the core sensing component.

8. The multimodal visual communication display device according to claim 2, characterized in that, The image acquisition device uses a binocular camera to recognize the user's three-dimensional gestures through a stereo vision algorithm. The recognizable gestures include clicking, swiping, zooming, and rotating, with a recognition accuracy of ≥95%. The gaze tracking sensor collects the coordinates of the user's eye feature points and calculates the focus coordinates of the user's gaze on the main display panel, with a positioning error of ≤5mm.

9. The multimodal visual communication display device according to claim 1, characterized in that, It also includes a wireless communication module, which is connected to the core control module and supports Wi-Fi, Bluetooth and NFC communication protocols, enabling multimodal data interaction with external terminals and remote control command reception.

10. The multimodal visual communication display device according to claim 1, characterized in that, The multimodal display execution module also includes a haptic feedback component, which is connected to the core control module. When the user triggers an interactive operation through gestures, the haptic feedback component outputs vibration feedback of different intensities according to the type of interaction. The vibration intensity is adjustable in 3-5 levels.