Man-machine interaction system combined with multifunctional super-screen projector

The human-computer interaction system, which combines a multi-functional ultra-screen projector with multiple modules and recognition technologies, solves the problem that human-computer interaction systems in multi-user environments cannot efficiently recognize and respond to diverse inputs, and achieves personalized user experience and system management optimization.

CN121501136APending Publication Date: 2026-02-10HANGZHOU ZHENWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511629802.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-08
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing human-computer interaction systems cannot efficiently support simultaneous operation by multiple users in a shared environment, and have weak recognition capabilities for non-touch or non-voice commands, resulting in a poor user experience, especially in noisy environments or when multiple people are using the system.

Method used

The human-computer interaction system, which combines a multi-functional ultra-screen projector, integrates a projection module, an interaction module, a user permission management module, a data acquisition module, a multi-screen collaboration module, an intelligent recognition module, a virtual reality module, a data analysis module, a remote control module, an environmental adaptive adjustment module, and an emotional interaction engine. Through touch, gesture, and voice recognition units, combined with deep learning models and a distributed computing framework, it achieves diverse input responses and optimizes the user experience through user permission management, emotional interaction, and environmental adaptive adjustment.

Benefits of technology

It improves the personalization of user interaction, enhances the convenience and security of interaction in multi-user environments, and improves user experience and the orderliness of system management by recognizing user emotional states through an emotional interaction engine and automatically adjusting the system feedback method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121501136A_ABST
    Figure CN121501136A_ABST
Patent Text Reader

Abstract

The invention discloses a man-machine interaction system of a multifunctional super-screen projector combination. The man-machine interaction system comprises a projection module, an interaction module, a user authority management module, a data acquisition module and the like. The projection module can perform multi-screen projection, and the interaction module integrates touch, gesture and voice recognition units and adopts a specific algorithm. The user permission management module manages users and permissions, the data acquisition module collects operation and state information, and the multi-screen collaboration module realizes multi-screen content interaction. The intelligent identification module responds to various inputs, the virtual reality module creates a virtual scene, the data analysis module optimizes the system, and the remote control module analyzes data and comprises an alarm sub-module. The environment adaptive module monitors environment adjustment projection parameters, the emotion interaction engine analyzes user state adjustment feedback, and the energy consumption optimization module dynamically adjusts power consumption. The user experience is improved, the interaction personalization is improved, the system is ensured to be safe and ordered, and the interaction convenience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of human-computer interaction applications, and in particular to a human-computer interaction system combining a multi-functional ultra-large screen projector. Background Technology

[0002] In the field of human-computer interaction, with the widespread adoption of smart devices, people are paying increasing attention to the interactive experience with these devices. Existing technologies achieve human-computer interaction by providing basic user interfaces and control functions, such as touchscreen operation and voice recognition. These technologies have improved the interactivity between users and devices to some extent, but they still have some limitations, especially in the interactive experience in complex scenarios, which needs further improvement.

[0003] Common human-computer interaction systems based on touch screens and voice recognition include touch screen interaction technology integrated into the operating systems of smartphones and tablets, as well as voice assistant services in smart speakers. However, this approach has limited applicability in multi-user shared environments, cannot efficiently support simultaneous operation by multiple users, and has weak recognition capabilities for non-touch or non-voice commands. This results in poor user experience in specific scenarios (such as noisy environments or when multiple people are using the system), failing to meet the working requirements of human-computer interaction applications. To address this, a multi-functional ultra-screen projector combined with a human-computer interaction system is proposed. Summary of the Invention

[0004] This invention provides the following technical solution: a human-computer interaction system combining a multi-functional ultra-large screen projector, comprising:

[0005] The system includes a projection module, an interaction module, a user permission management module, a data acquisition module, a multi-screen collaboration module, an intelligent recognition module, a virtual reality module, a data analysis module, a remote control module, an environmental adaptive adjustment module, an emotional interaction engine, and an energy consumption optimization module. The projection module is used to realize the multi-screen projection function, and the interaction module integrates touch, gesture, and voice recognition units.

[0006] The user permission management module is used to add, edit, and delete users, and to customize user permission levels. The data acquisition module is used to collect and store user operation records and device status information. The multi-screen collaboration module is used for content push, display, and data sharing between multiple screens.

[0007] The intelligent recognition module is used to quickly respond to diverse user inputs using image recognition, voice recognition, and gesture recognition algorithms. The virtual reality module is used to create virtual scenes and provide an immersive interactive experience. The data analysis module performs in-depth analysis of interactive data and extracts valuable information to optimize the system. The remote control module is used to perform in-depth analysis of interactive data and extract information. The environmental adaptive adjustment module is used to monitor ambient light, sound, and spatial layout in real time and dynamically adjust projection brightness, contrast, and interactive sensitivity.

[0008] The emotional interaction engine is used to analyze the user's facial expressions, voice tone and operating habits, identify the user's emotional state, and adjust the system feedback method. The energy consumption optimization module is used to dynamically adjust the system power consumption.

[0009] Preferably, the touch, gesture, and voice recognition units within the interaction module employ deep learning model optimization algorithms and a distributed computing framework.

[0010] Preferably, when adding a new user, the user permission management module, in addition to entering the username and password settings, simultaneously collects the user's biometric information, including fingerprints and facial features. The biometric information is bound to user permissions for higher-level security permission verification. The user permission management module is equipped with a role-based permission allocation function, and the permission roles include administrator, ordinary user, and visitor permissions.

[0011] Preferably, the multi-screen collaboration module has an internal HDMI interface, an internal wireless transmission module, and uses an SSL / TLS encryption protocol.

[0012] Preferably, the analysis methods of the data analysis module include data mining and machine learning algorithms. The data mining method is used to discover potential patterns in user operation behavior, that is, whether there is a periodic pattern in the frequency of user use of different module functions in different time periods. The machine learning algorithm is used to predict user behavior and needs.

[0013] Preferably, the remote control module uses virtual private network technology for remote data transmission. The remote control module has an alarm submodule installed inside. When the alarm submodule detects an abnormal situation in the system, it sends alarm information to the terminal via SMS.

[0014] Preferably, the environmental adaptive adjustment module includes a light sensor and a sound sensor. The light sensor is used to detect changes in light intensity, automatically increasing projection brightness in a bright light environment and reducing blue light output in a dark light environment. The sound sensor is used to detect ambient noise around the device.

[0015] Preferably, the projection module has an AI real-time translation and subtitle generation module submodule installed inside. The projection module adopts dynamic projection mapping technology, which is used to make the projected image automatically adapt to complex curved surfaces or non-planar display environments.

[0016] Preferably, the energy consumption optimization module dynamically adjusts the system power consumption through AI algorithms, so that it automatically enters a low-power standby mode when no one is using it.

[0017] Preferably, when the emotional interaction engine detects user fatigue and confusion, it automatically simplifies the interface and provides voice guidance; when the user shows excitement and focus, the emotional interaction engine simultaneously enhances the dynamic effects of the interaction.

[0018] In summary, compared with the prior art, the present invention provides a human-computer interaction system combining a multi-functional ultra-large screen projector, which has the following beneficial effects:

[0019] 1. This invention integrates touch, gesture, and voice recognition units within the interaction module, and employs deep learning model optimization algorithms and a distributed computing framework. This allows users to flexibly choose interaction methods according to their needs and usage scenarios, improving the user experience. Furthermore, the added emotional interaction engine can analyze the user's facial expressions, voice tone, and operating habits to identify the user's emotional state. When user fatigue or confusion is detected, the interface is automatically simplified and voice guidance is provided. When the user shows excitement or focus, the interactive dynamic effects are enhanced simultaneously, thereby improving the personalization of the user's interaction with the system.

[0020] 2. This invention, through the addition of a user permission management module, enables role-based permission allocation, including different permission roles such as administrator, ordinary user, and visitor permissions. This allows users with different roles to access corresponding functions and resources according to their needs and responsibilities. Administrators can manage the entire system, ordinary users can perform routine operations, and visitor permissions can restrict their operational scope. This not only facilitates use by different users but also ensures system security and orderly management. Furthermore, the added multi-screen collaboration module enables content push, display, and data sharing between multiple screens, improving the convenience of interaction. Attached Figure Description

[0021] Figure 1 This is a schematic diagram of the system structure of the present invention.

[0022] Figure 2 This is a schematic diagram of the interactive module structure of the present invention. Detailed Implementation

[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0024] Please see Figure 1 The present invention provides a technical solution, a human-computer interaction system for a multi-functional ultra-screen projector combination, comprising:

[0025] The system includes a projection module, an interaction module, a user permission management module, a data acquisition module, a multi-screen collaboration module, an intelligent recognition module, a virtual reality module, a data analysis module, a remote control module, an environmental adaptive adjustment module, an emotional interaction engine, and an energy consumption optimization module. The projection module is used to realize the multi-screen projection function. The projection module has an AI real-time translation and subtitle generation module sub-module installed inside. The projection module adopts dynamic projection mapping technology, which is used to make the projected image automatically adapt to complex curved surfaces or non-planar display environments. The specific process of the above steps is as follows:

[0026] Multi-screen projection functionality works as follows: First, the system receives video or image signals from external devices (such as computers, video players, etc.). These signals contain information about the content to be projected. The projection module identifies and parses the input signals, determining parameters such as signal format and resolution. The projection module also detects the currently connected screen devices, including the number and type of screens (e.g., flat screens, curved screens), resolution, and relative positions. This step determines how to distribute the input signals across the screens to achieve multi-screen projection. Based on the multi-screen configuration detection results, the projection module distributes the input video or image signals. For multiple flat screens, signals may be distributed according to a pre-defined layout (e.g., splicing, splitting). For non-flat screens, special processing using dynamic projection mapping technology is employed (described in detail in subsequent steps). Finally, the projection module projects the distributed signals onto the corresponding screens, completing the multi-screen projection function.

[0027] The AI ​​real-time translation and subtitle generation module works as follows: After the projection module receives the video or audio content to be projected, the AI ​​real-time translation and subtitle generation module first identifies the content. For video content, it extracts the speech information; for audio content, it directly acquires the speech signal. Simultaneously, it also identifies the text in the video content (such as subtitles and logos on the screen). After identifying the speech and text content, the module performs language detection to determine the original language type. Then, based on the user's settings or the system's default target language, it uses a pre-trained translation model to translate the content in real time. After translation, it generates corresponding subtitles based on the translation results. For video content, the subtitles are set according to a certain format (such as font, color, and position) and then blended with the original video screen to ensure that the subtitle display is clear, aesthetically pleasing, and does not affect the original screen content; for audio content, the subtitles are displayed separately on the projected screen, synchronized with the audio playback.

[0028] The dynamic projection mapping technology enables the projected image to automatically adapt to complex curved or non-planar display environments: When the projection module starts, it uses built-in sensors (such as depth sensors) to scan the target curved or non-planar environment. By scanning, it acquires the 3D shape information of the environment and then constructs a 3D model of the target display environment. This model includes information such as the shape, size, and surface material of the environment. Based on the constructed 3D model, the projection module corrects the original image to be projected. During the correction process, it considers the shape and angle changes of the curved or non-planar surface, performing stretching, distortion, and other transformations on the image to ensure that the image projected onto the target environment maintains the correct proportions and shape. Simultaneously, it calculates the mapping relationship between the image and the 3D model, determining the accurate projection position of each pixel on the target environment. Based on the calculated mapping relationship, the projection module adjusts the parameters of the projection lens (such as focal length and angle) to project the corrected image onto the complex curved or non-planar display environment. During the projection process, the projection is continuously optimized based on changes in the environment (such as minor surface deformations, light reflections, etc.). For example, parameters such as brightness and contrast are adjusted to ensure the visibility and quality of the projected image in complex environments.

[0029] Please see Figure 2 The interaction module integrates touch, gesture, and voice recognition units. The touch, gesture, and voice recognition units within the interaction module employ deep learning model optimization algorithms and a distributed computing framework. The specific implementation process of the interaction module is as follows:

[0030] Touch Recognition Unit Process: When a user touches an interactive device (such as a touchscreen), the touch sensor collects the signal generated by the touch operation. This signal contains information such as the coordinates of the touch point, touch pressure, and touch duration. The collected touch signal is sent to the preprocessing module of the touch recognition unit. Here, the signal undergoes preliminary processing, including noise removal and coordinate calibration. For example, the original coordinates of the touch point are converted into standard coordinates corresponding to the screen display area to ensure the accuracy of the touch operation. The preprocessed touch signal is then sent to the feature extraction module based on a deep learning model optimization algorithm. The deep learning model (such as a convolutional neural network) extracts features from the touch signal. These features may include the shape of the touch trajectory, the speed of the touch operation, and the pattern of touch pressure changes. These features reflect the user's touch operation intent and behavior pattern. The extracted feature data is distributed to multiple computing nodes in a distributed computing framework. On these computing nodes, a pre-trained classification model is used to identify the touch features. For example, it determines whether the touch operation is a click, swipe, or long press. Distributed computing frameworks can improve computing efficiency and speed up recognition. The recognition results are sent back to the interaction module, which then performs corresponding operations based on the recognition results, such as opening applications, switching pages, or adjusting settings.

[0031] Gesture recognition unit process: When a user performs a gesture, a camera or other image acquisition device captures an image sequence containing the gesture. First, the acquired images undergo preprocessing, including image filtering to remove noise, image enhancement to improve gesture clarity, and image normalization to adjust parameters such as image size and color to suit subsequent processing standards. In the preprocessed image, object detection technology based on deep learning model optimization algorithms is used to locate the region containing the gesture. This process identifies gesture-related parts of the image, eliminates irrelevant background information, and improves the accuracy and efficiency of subsequent processing. After locating the gesture region, a deep learning model (such as a recurrent neural network) extracts features from the image of the gesture region. These features may include the shape of the gesture, the trajectory of the gesture, and the relative positions of the fingers. These features are then converted into feature descriptors to represent the unique characteristics of the gesture. These feature descriptors are distributed to computing nodes in a distributed computing framework. At these nodes, the type of gesture, such as waving, liking, or grabbing, is identified by matching it with pre-trained gesture templates or gesture models. The distributed computing framework can process multiple gesture recognition tasks in parallel, improving recognition speed. The recognized gesture type is sent to the interaction module, which executes the corresponding operation according to the operation command of the gesture, such as controlling the switching of the projected image, zooming in and out, etc., and feeds back the operation result to the user.

[0032] Speech Recognition Unit Process: Voice acquisition devices such as microphones collect the user's voice signal. The voice signal is a continuous audio waveform containing the user's voice commands. The collected voice signal is sent to the preprocessing module of the speech recognition unit. Audio preprocessing operations are performed here, including background noise removal, volume adjustment, and audio framing. Audio framing divides the continuous voice signal into short frame sequences for subsequent processing. The preprocessed voice frames are then fed into a feature extraction module based on deep learning model optimization algorithms (such as deep neural networks). Acoustic features, such as Mel-frequency cepstral coefficients (MFCCs), are extracted from the voice frames. These features reflect the timbre, pitch, and duration of the speech. The extracted acoustic features are distributed to multiple computing nodes in a distributed computing framework. At these nodes, a pre-trained speech recognition model converts the acoustic features into corresponding text information. The distributed computing framework accelerates the speech recognition computation process and improves recognition efficiency. The recognized text information is then sent to the semantic understanding module, where natural language processing techniques are used to analyze the semantic content of the text. Based on the results of semantic understanding, the interaction module performs corresponding operations, such as searching for content and executing commands, and then returns the results to the user.

[0033] The user permission management module is used to add, edit, and delete users, and to customize user permission levels. When adding a user, in addition to entering the username and password, the user permission management module simultaneously collects the user's biometric information, including fingerprints and facial features. The biometric information is bound to user permissions for higher-level security permission verification. The user permission management module is equipped with a role-based permission allocation function, and permission roles include administrator, ordinary user, and visitor permissions.

[0034] The user addition process is as follows: When a new user needs to be added, the user first enters the user information entry interface of the user permission management module. The user or administrator enters a username, which should be unique to identify different users. Then, a password is set, which needs to meet certain security requirements (such as length and complexity). While entering the username and password, the system activates the biometric data acquisition device. For fingerprint acquisition, the user places their finger on the fingerprint sensor, which collects the fingerprint's ridge features. The fingerprint sensor scans the fingerprint multiple times to obtain a clear and complete fingerprint image, and then extracts key feature points such as ridges, valleys, and bifurcation points. For facial feature acquisition, the camera is turned on and captures the user's facial image. The system will prompt the user to keep their face in a suitable position and angle to ensure a clear facial image is captured. After the facial image is captured, image recognition technology is used to extract key facial features, such as the shape and position of the eyes, nose, and mouth, as well as facial contours. After completing the basic information (username and password) entry and biometric data acquisition, the administrator customizes the user's permission level according to the user's needs and role in the system. For example, if used internally by a company, permissions are determined based on employee positions and work requirements, binding the collected biometric information such as fingerprints and facial features with the newly defined user permissions. In the system's database, an association is established between biometric information and permission information. This association will be used for subsequent higher-level security permission verification. Based on the user's role in the system (administrator, regular user, or visitor), the user permission management module assigns corresponding function and resource access permissions. For administrators, full system access is granted, including adding, editing, and deleting users, managing system settings, and viewing and operating all modules. For regular users, permissions for routine operations are granted, such as using the projection module for projection and using the interactive module for normal interaction, but access to system management-related functions is restricted. For visitors, access to some public system resources is only allowed, such as viewing specific projection content and performing limited interactive operations, with restricted access to sensitive data and important system functions.

[0035] Access Control Process: When a user logs into the system, they first enter their username and password. The system compares the entered username and password with the information stored in the database. If the username and password match, the user proceeds to the next step; otherwise, login is denied and an error message is displayed. After successful username and password verification, the system initiates biometric verification. For fingerprint verification, the user places their finger on the fingerprint sensor again, and the system extracts the fingerprint features and compares them with previously bound fingerprint feature information. For facial verification, the camera captures the user's facial image, extracts facial features, and compares them with bound facial feature information. If biometric verification is successful, the user successfully logs in and accesses the corresponding functions and resources according to their authorized role; if biometric verification fails, login will be denied even if the username and password are correct, ensuring a higher level of security.

[0036] The data acquisition module is used to collect and store user operation records and device status information. The workflow of the above functions is as follows:

[0037] User Operation Record Collection: The data acquisition module continuously monitors various modules in the human-computer interaction system (such as the projection module and interaction module) to capture user operation events. For the projection module, it monitors user settings operations for multi-screen projection, such as screen selection and projection mode switching. In the interaction module, it monitors user interactions via touch, gestures, and voice recognition units, such as the location of touch operations, the type of gesture, and the content of voice commands. When a user operation event is detected, the data acquisition module immediately records the time of the operation. The time stamp is accurate to the specific moment for subsequent analysis of the order and frequency of operations. Time information provided by the system clock is used and recorded according to a specific time format (such as year, month, day, hour, minute, second) to determine the source of the operation, i.e., to determine which user initiated the operation. If multiple users are logged in, the module interacts with the user permission management module to obtain the identification information of the current user (such as username or user ID). Simultaneously, it identifies which device the operation was performed on. If multiple devices are connected (such as multiple terminal devices connected to the projector combination system), the device identification information (such as device number or device name) is recorded, along with detailed information about the operation. For example, in the interaction module, for touch operations, the system records parameters such as touch coordinates, touch duration, and touch pressure; for gesture operations, it records the gesture type (e.g., waving, liking) and the start and end positions of the gesture; for voice operations, it records the text content of the voice command and the acoustic characteristics of the voice (e.g., volume, tone). In the projection module, it records the adjustments to the projection screen's settings parameters (e.g., resolution, brightness).

[0038] Device Status Information Acquisition: The data acquisition module monitors the operational status of various device modules in the system (such as the projection module, interaction module, multi-screen collaboration module, etc.). For the projection module, it monitors the lifespan of the projection lamp, heat dissipation, and the status of the projection lens; in the interaction module, it monitors the sensitivity of the touch screen, the microphone status of the voice recognition unit, and the camera status of the gesture recognition unit; for the multi-screen collaboration module, it monitors the connection status of the HDMI interface and the signal strength of the wireless transmission module, and interacts with the environmental adaptive adjustment module to obtain environmental sensor data related to the device status. For example, it obtains ambient light intensity information from the light sensor, as light intensity may affect the brightness settings of the projection module and the display effect of the interaction module; it obtains the ambient noise level from the sound sensor, as noise level may affect the accuracy of the voice recognition unit, etc. Similar to user operation record collection, it records the time of device status information collection. It uses time information provided by the system clock and records it according to a specific time format (such as year, month, day, hour, minute, second) to analyze changes in device status over time.

[0039] Data storage: The collected user operation records and device status information are organized. Operation records are categorized by operation time, source, and details; device status information is categorized by device module, sensor type, and status parameters. This information is then converted into a suitable storage format, such as structured data (e.g., JSON or XML), to facilitate subsequent querying and analysis.

[0040] The multi-screen collaboration module is used for content push, display and data sharing between multiple screens. The multi-screen collaboration module has an HDMI interface installed inside, a wireless transmission module integrated inside, and uses SSL / TLS encryption protocol inside.

[0041] The intelligent recognition module utilizes image recognition, voice recognition, and gesture recognition algorithms to quickly respond to diverse user inputs. The virtual reality module creates virtual scenes to provide an immersive interactive experience. The data analysis module performs in-depth analysis of interactive data to extract valuable information for system optimization. The data analysis module's analysis methods include data mining and machine learning algorithms. Data mining is used to discover potential patterns in user behavior, i.e., whether there are periodic patterns in the frequency of user use of different module functions at different times. Machine learning algorithms are used to predict user behavior and needs. The specific process of the above functions is as follows:

[0042] First, data collection and integration are performed: The data collection module gathers user interaction data within the human-computer interaction system, including operation records, operation times, and operation sequences in various modules (such as the projection module and interaction module), as well as device status information. This interaction data covers detailed records of user interactions with touch, gesture, and voice recognition units, and multi-screen operation data in the multi-screen collaboration module. The collected interaction data is then integrated, with a unified data format and duplicate or invalid data removed. For example, timestamps in different formats are standardized, and similar operation data recorded in different modules are merged to form a complete interaction dataset, preparing for subsequent analysis.

[0043] Data mining analysis of potential patterns involves preprocessing the integrated interactive dataset, including data cleaning, removal of outliers (such as a large number of unreasonable operation records within a very short period), and handling of missing values ​​(such as using imputation or interpolation methods to handle missing operation records). Feature engineering is then performed to extract features related to user behavior, such as the time periods and frequency of use for each module's functions. Data mining algorithms are then used to analyze the preprocessed data to discover potential patterns in user behavior. Specifically, for the frequency of user use of different module functions at different time periods, statistical analysis and association rule mining are used to determine whether periodic patterns exist. For example, analyzing whether users use the multi-screen projection function of the projection module more frequently during specific time periods each day (such as 7-9 pm), or whether they tend to use the virtual reality module more often on certain days of the week.

[0044] Machine learning algorithms predict user behavior and needs: Based on the analysis objective (predicting user behavior and needs), a suitable machine learning algorithm is selected, such as decision trees or neural networks. The chosen machine learning model is trained using historical interaction data (a dataset containing labeled user behavior and needs). During training, the model's parameters are adjusted to accurately fit the data, meaning it can accurately predict user behavior and needs based on the characteristics of the input interaction data. The interaction data to be predicted (new data after preprocessing and feature extraction) is then input into the trained machine learning model. The model predicts the user's next behavior (e.g., whether to switch to a specific module function) and needs (e.g., whether to adjust projection brightness) based on the feature information in the data.

[0045] Information Extraction and System Optimization: By combining latent patterns discovered through data mining with user behavior and demand predictions from machine learning algorithms, valuable information for system optimization is extracted. For example, if it is discovered that users consistently need to adjust the projector brightness in specific scenarios, this is valuable information. This extracted valuable information is fed back to relevant modules to optimize the system. For instance, the environment adaptive adjustment module is notified to adjust the projector brightness in advance based on user habits, or the interaction module is notified to optimize the interaction flow of specific functions to improve the user experience.

[0046] Furthermore, after extracting and optimizing information: the DoWhy framework is used to distinguish the causal relationship between user behavior and environmental changes (such as distinguishing between misoperation caused by insufficient brightness and user operation errors), the model is trained locally on the desensitized user behavior data, only the aggregated gradient update parameters are uploaded to the cloud, the newly generated interaction strategy is released in a gray-scale manner, the system efficiency indicators of the new and old versions are compared by grouping by the last digit of the user ID, and when frequent abnormal operations are detected, the device logs of the past 24 hours are automatically traced back, and the fault module with the highest probability is located through a Bayesian network;

[0047] The remote control module is used to perform in-depth analysis of interactive data and extract information. The remote data transmission of the remote control module adopts virtual private network technology. The remote control module has an alarm submodule installed inside. When the alarm submodule detects an abnormal situation in the system, it sends alarm information to the terminal via SMS. The specific process of the above functions is as follows:

[0048] First, data acquisition is performed: the remote control module obtains interactive data from the data acquisition module. This data includes the user's operation records in various modules of the multi-functional ultra-screen projector combination human-computer interaction system (such as the projection module, interaction module, etc.), including information such as the type of operation, the time of operation, the order of operation, and the status information of the device;

[0049] Next, data transmission occurs: Before remote data transmission, the remote control module initiates Virtual Private Network (VPN) technology to establish a secure connection with the target device (such as a server or other remote terminal). The VPN connection ensures security and privacy during data transmission, and the acquired interactive data is sent to the remote analysis and processing unit via the established VPN connection. During transmission, the data is processed according to predetermined encryption and encapsulation protocols to ensure data integrity and accuracy.

[0050] In-depth analysis and information extraction: After receiving the interactive data, the remote analysis and processing unit first preprocesses the data. This includes data cleaning, removing outliers (such as unreasonable operation records or incorrect device status values) and handling missing values ​​(using appropriate imputation or interpolation methods), extracting data features such as those related to user behavior, including operation frequency and time periods, to prepare for subsequent analysis, and using specific analysis algorithms (similar to methods in the data analysis module or other algorithms suitable for remote analysis) to conduct in-depth analysis of the preprocessed data. For example, analyzing user operation patterns and trends to uncover underlying rules in user behavior, and extracting valuable information for system management, optimization, or monitoring through the analysis of interactive data. This information may include user preferences for specific functions, potential system failure risks, etc.

[0051] Anomaly Detection and Alarm: The alarm submodule within the remote control module performs anomaly detection on the analyzed data. It determines whether the system is experiencing abnormal conditions based on predefined rules and models. For example, a sudden and abnormal increase or decrease in user operation frequency, or unreasonable values ​​in device status, could be considered anomalies. When the alarm submodule detects an anomaly, it generates an alarm message containing detailed anomaly information (such as anomaly type, occurrence time, and potentially affected modules). This alarm message is then sent via an SMS module to a pre-defined terminal (such as the administrator's mobile phone) so that relevant personnel can take timely measures to handle the anomaly.

[0052] The environmental adaptive adjustment module is used to monitor the surrounding light, sound and spatial layout in real time, and dynamically adjust the projection brightness, contrast and interactive sensitivity. The environmental adaptive adjustment module includes a light sensor and a sound sensor. The light sensor is used to detect changes in light intensity, automatically increase the projection brightness in a strong light environment, and reduce blue light output in a dark light environment. The sound sensor is used to detect the ambient noise around the device. The specific process of the above functions is as follows.

[0053] Light monitoring and projection brightness / blue light output adjustment: The light sensor continuously detects the light intensity around the device. It converts the light intensity into an electrical signal, the value of which corresponds to the light strength. For example, the stronger the light, the larger the electrical signal value; the weaker the light, the smaller the electrical signal value. When the electrical signal value detected by the light sensor exceeds a preset strong light threshold, it is determined to be a strong light environment. The environment adaptive adjustment module sends a command to the projection module to increase the projection brightness. After receiving the command, the projection module increases the brightness of the projected image through its internal brightness adjustment mechanism (such as adjusting the lamp power) to ensure that the projected image remains clearly visible in strong light environments. When the electrical signal value detected by the light sensor is lower than a preset dark light threshold, it is determined to be a dark light environment. The environment adaptive adjustment module sends a command to the projection module to reduce blue light output. The projection module adjusts its display settings according to the command to reduce blue light components. This reduces the stimulation of the user's eyes by blue light in dark environments, and may also adjust the proportion of other colors to maintain the color balance of the image.

[0054] Sound Monitoring and Interaction Sensitivity Adjustment: The sound sensor detects ambient noise around the device in real time. The sound sensor converts sound signals into digital signals, which reflect the volume, frequency, and other characteristics of the ambient noise. The environmental adaptive adjustment module analyzes the digital signals collected by the sound sensor according to a preset algorithm. If the ambient noise is high, it may interfere with the accuracy of voice recognition or other interactive operations. When the ambient noise exceeds a certain threshold, the environmental adaptive adjustment module determines that the interaction sensitivity needs to be adjusted. For example, for voice interaction, it may be necessary to lower the voice recognition sensitivity threshold to reduce misrecognition; for touch or gesture interaction, it may be necessary to adjust the sensitivity parameters of the touch screen or gesture recognition to accommodate potential misoperations in noisy environments. The environmental adaptive adjustment module sends an instruction to the interaction module to adjust the interaction sensitivity. After receiving the instruction, the interaction module adjusts the sensitivity settings of the touch, gesture, and voice recognition units according to the parameters in the instruction. For example, adjusting the filtering parameters or recognition threshold of the voice signal in the voice recognition unit.

[0055] Spatial Layout Monitoring and Projection Adjustment: The environmental adaptive adjustment module uses sensors, such as distance sensors, to acquire information about the distance between the projection device and surrounding obstacles, as well as the shape and size of the projection plane. Based on the collected spatial layout information, the environmental adaptive adjustment module analyzes how to adjust the projection settings to adapt to the spatial layout. For example, if the projection plane is not a standard rectangle, it may be necessary to adjust the parameters of the projection mapping technology to ensure that the projected image can correctly adapt to the non-standard projection plane. The environmental adaptive adjustment module sends adjustment instructions based on the spatial layout to the projection module. The projection module adjusts parameters such as the mapping method and aspect ratio of the projected image according to the instructions. For example, dynamic projection mapping technology can be used to automatically adapt the projected image to complex curved surfaces or non-planar display environments.

[0056] After dynamically adjusting the projection parameters: the system monitors changes in light refraction in the physical space in real time through tiny calibration patterns at the edges of the projected image (refreshed twice per second). When the internal temperature of the projector exceeds a threshold, it automatically reduces the brightness output and triggers the cooling fan to speed up to prevent image distortion caused by overheating. At the same time, it records user interventions in the automatic adjustment (such as manually dimming the brightness), optimizes the weights of subsequent environmental parameter calculations through reinforcement learning algorithms, and communicates with the intelligent lighting system. When the meeting room curtains are detected to be automatically closed, the system simultaneously increases the projection contrast to cope with sudden changes in ambient light.

[0057] The emotional interaction engine is used to analyze users' facial expressions, voice tone and operating habits, identify users' emotional states, and adjust the system's feedback methods. When the emotional interaction engine detects user fatigue and confusion, it automatically simplifies the interface and provides voice guidance. When users show excitement and focus, the emotional interaction engine simultaneously enhances the interactive dynamic effects. The specific process of the above functions is as follows:

[0058] First, data acquisition is performed: The image recognition algorithm in the intelligent recognition module and the camera capture images of the user's face. The camera acquires image frames of the user's face at a certain frame rate (e.g., 30 frames per second). The acquired facial images contain various facial features, such as the shape information of the eyes, eyebrows, and mouth. The user's voice input is then acquired through the voice recognition unit in the interaction module. While acquiring the voice signal, the acoustic characteristics of the voice are recorded, including pitch (e.g., changes in pitch), speech rate, and volume. The data acquisition module collects the user's operation records, including the user's actions in various modules (e.g., the projection module, the interaction module, etc.). For example, information such as the frequency of the user's use of different functions, the order of operations, and the time intervals between operations.

[0059] Emotional State Recognition: The emotional interaction engine utilizes a pre-trained facial expression recognition model to analyze captured facial images. This model, trained on a large amount of facial expression data, can identify different facial expression patterns, such as happiness, sadness, fatigue, and confusion. It analyzes the position and shape changes of various facial feature points, such as whether the eyes are squinting, eyebrows are furrowed, or the mouth is downturned, to determine the emotional state represented by the user's facial expression. Secondly, it analyzes the captured voice tone data. For example, a slow speaking speed and low tone may be associated with fatigue or confusion; while a fast speaking speed and high tone may be associated with excitement or focus. By comparing with predefined voice tone patterns, the system determines the user's emotional state. Based on user operation habit data, it analyzes the user's operation patterns. For example, frequent pauses and repetitive operations may indicate confusion; fast and smooth operations may indicate focus or excitement. Through statistical analysis and pattern recognition techniques, the system infers the user's emotional state from operation habit data, and comprehensively judges the results obtained from facial expression, voice tone, and operation habit analysis. For example, if the facial expression shows fatigue, the tone of voice is slow, and there are many pauses in the operation, then it can be judged that the user is in a state of fatigue and confusion; if the facial expression shows excitement, the tone of voice is light, and the operation is fast and smooth, then it can be judged that the user is in a state of excitement and focus.

[0060] System feedback method adjustment: The system feedback method has been adjusted so that the interaction module automatically simplifies the interface based on instructions, such as reducing the amount of information displayed on the interface and simplifying the operation process (e.g., hiding some infrequently used function buttons). Simultaneously, the emotional interaction engine triggers the voice guidance function, providing operation guidance to the user through voice prompts, such as "You can click here to proceed to the next step." When the emotional interaction engine determines that the user is in an excited and focused state, it sends instructions to the interaction module. Upon receiving the instructions, the interaction module simultaneously enhances the dynamic effects of the interaction. For example, it adds animation effects to the interactive interface, improves the feedback sensitivity of touch operations (e.g., enhances touch vibration feedback), and optimizes the speech synthesis effect of voice interaction (e.g., using a more vivid voice tone).

[0061] When user fatigue and confusion are detected:

[0062] Automatic Interface Simplification: The emotional interaction engine sends simplification instructions to the interface management module. Upon receiving the instructions, the interface management module first analyzes the layout of the current interface to identify non-core functional areas and elements, such as pop-up ads and secondary setting options. Then, by calling the interface rendering submodule, these non-core elements are hidden or shrunk, while the size and spacing of core operation buttons are increased to ensure that users can more easily identify and click them.

[0063] Provides voice guidance: The emotional interaction engine transmits the user's current task and points of confusion to the voice interaction module. Based on a preset voice guidance template, the voice interaction module generates personalized voice prompts tailored to the user's specific situation. For example, if a user shows confusion while setting projection parameters, the voice interaction module will generate clear and detailed voice guidance such as, "You are currently setting projection parameters. To adjust brightness, please click the brightness adjustment button; to adjust contrast, please click the contrast adjustment button." The voice prompts are then played back to the user via an audio output device.

[0064] Subsequent monitoring and adjustments: The emotional interaction engine continuously collects users' facial expressions and voice feedback through the camera and microphone, while also recording user actions. If it detects that the user's operation speed increases, the error rate decreases, and their facial expressions become more relaxed, it indicates that the user's fatigue and confusion have been alleviated. At this point, the emotional interaction engine sends a gradual recovery command to the interface management module. The interface management module, according to a preset recovery strategy, gradually displays previously hidden non-core elements and adjusts the size and spacing of core buttons back to their original state. Simultaneously, the voice interaction module reduces the frequency of voice guidance, only providing appropriate prompts when the user has not operated for an extended period or when obvious errors occur.

[0065] If the user's condition does not improve: If the user's fatigue and confusion persist or worsen, the emotional interaction engine further simplifies the interface, retaining only the most essential operation buttons and key information displays. Simultaneously, the voice interaction module enhances the detail of voice guidance, using a slower, clearer speech rate and adding interactive elements, such as asking the user if they understand the current operation steps, and providing targeted guidance based on the user's voice response. Furthermore, the emotional interaction engine will consider introducing other auxiliary interaction methods, such as sending a start command to the gesture recognition module, allowing users to complete some basic operations through simple gestures, reducing the operational burden.

[0066] When user excitement and focus are detected:

[0067] Synchronous Enhancement of Interactive Dynamics: The emotional interaction engine sends enhancement commands to the animation effects management module. Upon receiving the commands, the animation effects management module analyzes the dynamic effects of elements in the current interactive interface, identifying elements suitable for adding dynamic effects, such as buttons, icons, and text. Then, based on a pre-set dynamic effects library, it adds appropriate dynamic effects to these elements, such as scaling animations for buttons, rotation animations for icons, and gradient animations for text. Simultaneously, it adjusts the parameters of the dynamic effects, such as animation speed and amplitude, to make the dynamic effects more pronounced and vivid.

[0068] Subsequent Monitoring and Adjustment: The emotional interaction engine continuously monitors the user's emotional state and interactive behavior. If the user remains excited and focused, the animation effect management module maintains or further enhances the dynamic effects, such as increasing the complexity and diversity of the animations and introducing more interactive feedback animations. For example, when the user clicks a button, the button not only scales but also pops up a small icon to indicate that the operation was successful. If the user's emotional state changes, such as becoming calm or fatigued, the animation effect management module adjusts the intensity and frequency of the dynamic effects in a timely manner, reducing the complexity and amplitude of the animations to avoid overstimulating the user or causing visual fatigue.

[0069] The energy consumption optimization module is used to dynamically adjust the system power consumption. The energy consumption optimization module dynamically adjusts the system power consumption through AI algorithms, so that it automatically enters a low-power standby mode when no one is using it. The specific process of the above function is as follows:

[0070] First, data collection and analysis are performed: The data collection module gathers user operation records for various modules of the human-computer interaction system (such as the projection module and the interaction module), including the operation time, operation type (such as turning on the projection, switching functions, etc.), and operation frequency. Simultaneously, it collects the operating status information of each module, such as the projection duration of the projection module and the frequency of interactive activities in the interaction module. The energy consumption optimization module uses AI algorithms to analyze the collected data. The AI ​​algorithm is pre-trained based on a large amount of historical data, including system power consumption under different user operation modes. By analyzing user operation data and module operating status, it predicts the system's usage in the future and determines whether it is likely to enter an unused state. For example, if no user operation is detected on any module within a certain period (e.g., 10 minutes), and there is no ongoing task (such as a playing video or a running interactive program), it is preliminarily determined that the system may be about to enter an unused state.

[0071] Dynamic power consumption adjustment: First, power consumption is adjusted during normal use. For infrequently used functional modules or sub-modules currently idle, their power consumption is reduced. For example, if the virtual reality module is not used for a period of time, the power optimization module can reduce its power supply, putting it into a low-power operating mode (e.g., reducing processor frequency, turning off unnecessary sensors), and rationally allocate power resources according to current operational needs. For example, when the projection module is performing high-resolution multi-screen projection, the power optimization module ensures sufficient power supply to the projection module, while appropriately reducing the power consumption of other non-critical modules (such as some auxiliary sensors in the environmental adaptive adjustment module). For adjustments when approaching a state of near-unused operation, when data analysis indicates that the system is approaching a state of near-unused operation, the power optimization module further reduces the overall system power consumption. First, it notifies each module to gradually reduce unnecessary activities. For example, it notifies the interaction module to stop some background interaction monitoring (such as continuous monitoring of gesture recognition), reducing its power consumption. The system notifies the projection module to reduce brightness or switch to a low-power display mode (such as reducing resolution or refresh rate), and continuously monitors the system status. If no user operation is detected within a certain period of time (such as 2 minutes), and the system activity has been reduced to a minimum, it is determined to be in an unused state.

[0072] Entering low-power standby mode: When it is determined that no one is using the system, the power optimization module sends a command to each module to enter low-power standby mode. For the projection module, the projection lamp is turned off or switched to a standby screen display with extremely low power consumption. The operation of advanced functions such as dynamic projection mapping technology is stopped, and only the most basic wake-up function monitoring is retained. The interaction module stops the operation of touch, gesture and voice recognition units, and only retains a small amount of power for detecting wake-up signals (such as detecting specific touch wake-up gestures or voice wake-up commands). Other modules such as the intelligent recognition module and data analysis module also stop most functions, and only retain minimal power consumption to maintain the basic state of the system, waiting for the wake-up command.

[0073] Low-power standby mode maintenance: In low-power standby mode, the power optimization module continuously monitors the system's wake-up signals. For example, it detects whether there are user touch wake-up gestures, voice wake-up commands, or wake-up signals from external devices (such as remote control signals). Once a wake-up signal is detected, the power optimization module immediately notifies each module to return from low-power standby mode to normal operating state and reinitializes each module according to the normal startup process to ensure that the system can quickly respond to subsequent user operations.

[0074] After entering low-power standby mode:

[0075] Tiered hibernation strategy: Functions are turned off in order of module importance—first, non-core virtual reality modules are paused, then the refresh rate of the projection module is reduced to 15Hz, and finally the power supply to the environmental sensors is cut off.

[0076] Heartbeat monitoring mechanism: The low-power Bluetooth module is retained to continuously listen for wake-up signals. When the phone's NFC is detected to be close or a specific voice command is given, a fast wake-up process is triggered within 300ms.

[0077] Data insulation: Maintain power supply to critical data in memory (such as user permission information and recent operation records) during standby to ensure a seamless transition after recovery;

[0078] Self-test cycle program: The main control chip is woken up every 2 hours to perform hardware status inspection. When an abnormality is found, a local alarm is issued through a buzzer and LED lights.

[0079] This solution integrates touch, gesture, and voice recognition units within the interaction module, employing deep learning model optimization algorithms and a distributed computing framework. This allows users to flexibly choose interaction methods based on their needs and usage scenarios, improving the user experience. Furthermore, the added emotional interaction engine analyzes users' facial expressions, voice tone, and operating habits to identify their emotional state. When user fatigue or confusion is detected, the interface is automatically simplified and voice guidance is provided. When users exhibit excitement or focus, the interactive dynamics are enhanced, thereby increasing the personalization of user interaction with the system.

[0080] This solution incorporates a user access control module, enabling role-based permission allocation. This includes different permission roles such as administrator, regular user, and visitor. Users with different roles can access corresponding functions and resources based on their needs and responsibilities. Administrators have comprehensive system control, regular users can perform routine operations, and visitor permissions restrict their scope of operation. This approach facilitates use by different users while ensuring system security and orderly management. Furthermore, the added multi-screen collaboration module enables content pushing, display, and data sharing across multiple screens, enhancing the convenience of interaction.

[0081] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0082] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A human-computer interaction system combining a multi-functional ultra-wide screen projector, characterized in that, include: The system includes a projection module, an interaction module, a user permission management module, a data acquisition module, a multi-screen collaboration module, an intelligent recognition module, a virtual reality module, a data analysis module, a remote control module, an environmental adaptive adjustment module, an emotional interaction engine, and an energy consumption optimization module. The projection module is used to realize the multi-screen projection function, and the interaction module integrates touch, gesture, and voice recognition units. The user permission management module is used to add, edit, and delete users, and to customize user permission levels. The data acquisition module is used to collect and store user operation records and device status information. The multi-screen collaboration module is used for content push, display, and data sharing between multiple screens. The intelligent recognition module is used to respond to user input using image recognition, sound recognition, and gesture recognition algorithms. The virtual reality module is used to create virtual scenes. The data analysis module is used to perform in-depth analysis of interactive data, extract information, and optimize the system. The remote control module is used to analyze interactive data and extract information. The environmental adaptive adjustment module is used to monitor light, sound, and spatial layout in real time, and dynamically adjust projection brightness, contrast, and interactive sensitivity. The emotional interaction engine is used to analyze the user's facial expressions, voice tone and operating habits, identify the user's emotional state, and adjust the system feedback method. The energy consumption optimization module is used to dynamically adjust the system power consumption.

2. The human-computer interaction system of a multi-functional ultra-wide screen projector combination according to claim 1, characterized in that: The touch, gesture, and voice recognition units within the interaction module employ deep learning model optimization algorithms and a distributed computing framework.

3. The human-computer interaction system of a multi-functional ultra-wide screen projector combination according to claim 1, characterized in that: When adding a new user, the user permission management module, in addition to entering the username and password, simultaneously collects the user's biometric information, including fingerprints and facial features. The biometric information is bound to user permissions for higher-level security permission verification. The user permission management module is equipped with a role-based permission allocation function, and the permission roles include administrator, ordinary user, and visitor permissions.

4. The human-computer interaction system of a multi-functional ultra-wide screen projector combination according to claim 1, characterized in that: The multi-screen collaboration module is equipped with an HDMI interface and integrates a wireless transmission module. The multi-screen collaboration module uses SSL / TLS encryption protocol.

5. The human-computer interaction system of a multi-functional ultra-wide screen projector combination according to claim 1, characterized in that: The data analysis module employs data mining and machine learning algorithms. The data mining method is used to discover potential patterns in user behavior, namely the periodic patterns of user frequency of using different module functions at different times. The machine learning algorithm is used to predict user behavior and needs.

6. The human-computer interaction system of a multi-functional ultra-wide screen projector combination according to claim 1, characterized in that: The remote control module uses virtual private network technology for remote data transmission. An alarm submodule is installed inside the remote control module. When the alarm submodule detects an abnormal situation in the system, it sends alarm information to the terminal via SMS.

7. The human-computer interaction system of a multi-functional ultra-wide screen projector combination according to claim 1, characterized in that: The environmental adaptive adjustment module includes a light sensor and a sound sensor. The light sensor is used to detect changes in light intensity, automatically increasing projection brightness in strong light environments and reducing blue light output in dark light environments. The sound sensor is used to detect ambient noise around the device.

8. The human-computer interaction system of a multi-functional ultra-wide screen projector combination according to claim 1, characterized in that: The projection module has an AI real-time translation and subtitle generation submodule installed inside. The projection module adopts dynamic projection mapping technology, which is used to make the projected image automatically adapt to complex curved surfaces and non-planar display environments.

9. The human-computer interaction system of a multi-functional ultra-wide screen projector combination according to claim 1, characterized in that: The energy consumption optimization module dynamically adjusts the system power consumption through AI algorithms, so that it automatically enters a low-power standby mode when no one is using it.

10. The human-computer interaction system of a multi-functional ultra-wide screen projector combination according to claim 1, characterized in that: When the emotional interaction engine detects user fatigue and confusion, it automatically simplifies the interface and provides voice guidance. When the user shows excitement and focus, the emotional interaction engine simultaneously enhances the dynamic effects of the interaction.