Human-computer interaction intelligent screen element boundary capable of achieving nature

The Smart Screen Metaverse system, designed with multiple modules in synergy, solves the problems of high cost, poor portability, and insufficient interaction capabilities of AR/VR devices. It enables natural human-computer interaction and multi-device collaboration, supports high-precision gesture and voice recognition, provides realistic spatial perception, and is suitable for multi-scenario applications.

CN121560424AInactive Publication Date: 2026-02-24宫贺辰
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511388636.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-26
Publication Date
2026-02-24
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing AR/VR devices are expensive, not portable, lack natural interaction capabilities, and cannot achieve multi-device collaboration and cross-scene expansion. Traditional display technologies lack stereoscopic capabilities, and existing projection technologies cannot achieve real spatial perception.

Method used

The Smart Screen Metaverse system, which adopts a multi-module collaborative design, includes an ultra-high-definition display chip, precision optical components, multi-modal sensing technology, edge computing architecture, and 5G network module. It achieves stereoscopic image display through Micro LED technology, integrates depth sensors, microphone arrays, and multi-sensor data fusion, supports high-precision interaction and real-time computing, builds a high-speed, low-latency network architecture, and realizes natural human-computer interaction.

Benefits of technology

Significantly reduces device costs, enables natural interaction and multi-device collaboration, supports high-precision gesture and voice recognition, provides realistic spatial awareness, meets the needs of multiple application scenarios, has a latency of less than 10ms, and is suitable for open environments such as homes, shopping malls, and exhibitions.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The invention discloses a man-machine interaction intelligent screen element boundary capable of achieving nature, and deep fusion of virtuality and reality is achieved through the ultra-high-definition stereo projection and multi-mode interaction technology. The system does not need a special head display device, generates a dynamic stereoscopic image by using a screen of an existing intelligent device, realizes natural interaction in combination with gestures, voice and physical touch, is suitable for the fields of games, education, social contact and the like, and provides an efficient and low-cost solution for the global universe.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of virtual reality and augmented reality technology, specifically relating to a smart screen metaverse that enables natural human-computer interaction. Background Technology

[0002] With the rapid development of virtual reality (VR), augmented reality (AR), and the metaverse concept, traditional display technologies and interaction methods are gradually revealing the following limitations: Existing AR / VR devices (such as headsets) rely on dedicated hardware, resulting in high costs and poor portability. For example, while silicon-based OLED microdisplays can provide high-resolution near-eye displays, their cost still accounts for nearly 50% of the total device cost, and they require complex optical modules, making the devices bulky; traditional touch and controller operations are difficult to meet the needs of natural interaction, while emerging holographic imaging technologies (such as the suspended 3D images developed by a Spanish team) support touch, but rely on rigid materials and complex calibration systems, and have not yet achieved commercialization; existing projection technologies (such as 2D screen projection) lack stereoscopic capabilities and cannot achieve realistic spatial perception, while high-resolution OLED displays (such as 12864 dot matrix screens) have vibrant colors, but can only display on a flat surface and cannot break through dimensional limitations; existing technical solutions mostly focus on single devices (such as mobile phones or VR headsets), lacking multi-device collaboration and cross-scene expansion capabilities, making it difficult to support the open ecosystem required by the "Earth metaverse". Summary of the Invention

[0003] The purpose of this invention is to provide a smart screen metaverse that enables natural human-computer interaction.

[0004] To achieve the above objectives, this invention provides a smart screen metaverse system enabling natural human-computer interaction. Its core feature lies in achieving an immersive interactive experience that blends the virtual and real worlds through multi-module collaboration. The system consists of four main functional modules: the projection module utilizes an ultra-high-definition display chip and precision optical components to convert the content of the smart device screen into an aerial stereoscopic image. The display chip employs Micro LED technology, with a pixel size of only 0.002 mm, a pixel density of 1600 PPI, and a contrast ratio of 1,500,000:1, maintaining a peak brightness of 1500 nits even under direct sunlight. The optical refraction component achieves image focusing through a three-layer Fresnel lens array. The first lens has a radius of curvature of 30 mm, the second 50 mm, and the third 70 mm. This three-layer lens combination converts a 2D image into a stereoscopic image with a diameter of 500 mm, while controlling the image distortion rate to within 0.5%.

[0005] Furthermore, the interaction module integrates multimodal sensing technology. The depth sensor employs Time-of-Flight (ToF) ranging technology, using 940nm infrared light pulses to achieve spatial ranging with a measurement accuracy of ±0.8 mm (at a distance of 1 meter), a frame rate of up to 180fps, and an effective recognition distance extended to 8 meters. It can capture the movement trajectory of human skeletal joints with a joint recognition accuracy of 1.5 mm. The microphone array consists of eight MEMS microphones arranged in a ring, employing beamforming algorithms and adaptive noise reduction technology to achieve far-field voice recognition within a 5-meter range. It supports multilingual command set parsing, including Chinese, English, and Japanese, with a voice command recognition accuracy of 98.2%. It also features echo cancellation to cope with complex acoustic environments. Through multi-sensor data fusion, this module supports millimeter-level limb motion capture and high-precision interactive response. The 0.5 Newton force generated by a user's arm movement can be accurately captured by the sensor and converted into a digital signal.

[0006] Furthermore, the computing module is designed based on an edge computing architecture, incorporating the NVIDIA Jetson Xavier NX computing platform and equipped with an AI acceleration chip boasting 21 TOPS of computing power. It utilizes TSMC's 12nm process technology, with power consumption controlled below 15 watts. The physics engine, based on rigid body dynamics and SPH fluid simulation algorithms, supports 5 million collision detection operations per second, capable of simulating the motion trajectories and interactions of 1000 objects in a virtual scene in real time. Through a dynamic rendering engine and AI inference acceleration technology, the system can complete user behavior analysis and scene response within 8 milliseconds, with AI inference latency controlled below 12 milliseconds, ensuring real-time performance and smoothness in the interaction process. This module, through hardware and software co-optimization, balances computational efficiency and energy efficiency, maintaining a temperature below 45 degrees Celsius under full load.

[0007] Furthermore, the network module constructs a high-speed, low-latency network architecture based on the 5G NR communication protocol (millimeter-wave band), employing 28GHz frequency band communication with an air interface latency as low as 7.2 milliseconds and a theoretical throughput of 10Gbps. It can support 256 nodes simultaneously online, with a scene update frequency of 120Hz, meeting the needs of high-concurrency multi-user collaboration. Through adaptive bitrate control technology and distributed network protocols, the system can dynamically adjust data transmission strategies according to network bandwidth, achieving stable transmission while ensuring image quality. This module constructs a decentralized network architecture through edge computing and cloud computing collaboration, supporting local data processing and cloud resource scheduling, with a data packet loss rate of less than 0.1% in multi-user collaboration scenarios.

[0008] Furthermore, the human-computer interaction method employs a phased collaborative processing mechanism to achieve a natural interaction flow: First, a dynamic compensation algorithm (based on temporal interpolation and motion estimation) eliminates projection aliasing effects, combined with multi-sampling anti-aliasing technology (SSAA+FXAA hybrid mode), improving image edge sharpness by 65% ​​and image clarity by 45% after dynamic compensation. Second, it integrates 26-DOF gesture recognition technology, supporting gesture trajectory capture with 0.05 mm precision and completing recognition within 400 milliseconds; voice commands support Natural Language Understanding (NLU), covering multilingual command set parsing and capable of parsing complex commands containing contextual semantics. Third, the ambient light sensor uses a wide dynamic range photodetector to monitor the light intensity range of 10-100,000 lux in real time, achieving stepless adjustment of projection brightness with an adjustment accuracy of 0.5% through PWM dimming technology.

[0009] Furthermore, the user behavior data acquisition module's depth sensor supports 26 degrees of freedom gesture recognition, covering fine movements such as finger flexion and extension, palm opening and closing, and arm swinging, with a recognition speed of less than 350ms response time; the microphone array supports natural language understanding (NLU), which can parse compound commands containing contextual semantics, supports command sets in seven languages ​​including English, Chinese, and Japanese, and achieves a dialect recognition accuracy of 92%; the ambient light sensor uses an indium gallium arsenide photodiode to monitor the light intensity range of 10-100,000 lux in real time, and achieves dynamic adjustment of projection brightness through a PID control algorithm with a response time of less than 200ms.

[0010] Furthermore, the ultra-high-definition image processing technology employs a dynamic compensation algorithm (based on temporal interpolation and motion estimation) to eliminate projection aliasing, combined with multi-sampling anti-aliasing technology (SSAA+FXAA hybrid mode), improving image edge sharpness by 70%. The dynamic compensation algorithm analyzes inter-frame motion vectors to predict pixel distribution and combines motion estimation to compensate for missing pixel information, effectively reducing motion blur in high-speed scenes. The anti-aliasing technology combines oversampling and blur filtering to eliminate high-frequency aliasing while maintaining image smoothness. This technology enables 4K image edge sharpness to achieve retina-level display effects, with dynamic scene rendering frame rates consistently above 120fps.

[0011] Furthermore, the virtual scene development framework provides native plugins for the Unity / Unreal Engine, supporting custom interaction and physics rules via RESTful APIs. Users can extend the physics engine's functionality using Python / JavaScript scripting languages, defining parameters such as collision detection logic, gravity parameters, and material properties. The development framework includes a built-in visual script editor that supports drag-and-drop logic arrangement, lowering the development threshold. The physics engine adopts the open-source PhysX engine architecture, supporting rigid / soft body hybrid simulation, and can simulate complex physical phenomena such as cloth deformation and fluid dynamics, with fluid simulation resolution reaching 1 millimeter level.

[0012] Furthermore, in terms of system compatibility and scalability, it is compatible with Android 10 and above, iOS 14 and above, and Linux Ubuntu 20.04 and above operating systems, achieving hardware interface unification through the HIDL / DBus driver adaptation layer. It supports hot-swappable peripheral access, including external sensors, controllers, and other devices, offering plug-and-play functionality without requiring a system restart. The driver adaptation layer adopts a modular design, allowing dynamic loading of new hardware drivers to adapt to future new interactive devices. This feature gives the system excellent backward compatibility and scalability, enabling rapid integration into the existing smart device ecosystem. It has passed multiple international certifications such as FCC, CE, and RoHS, and its electromagnetic compatibility reaches CLASS A level.

[0013] This invention provides a smart screen metaverse that enables natural human-computer interaction, and has the following beneficial effects:

[0014] (1) Use existing device screens such as mobile phones / tablets to replace dedicated head-mounted displays, and combine silicon-based OLED micro-display technology to significantly reduce equipment costs; through Fresnel lens arrays and dynamic compensation algorithms, screen content is converted into aerial stereoscopic images, eliminating the screen door effect and achieving ultra-high-definition and realistic spatial perception.

[0015] (2) It integrates an infrared camera (to capture gesture trajectory), a microphone array (for far-field speech recognition), and a depth sensor (for millimeter-level positioning), supporting multi-channel interaction such as gesture grasping, voice commands, and physical touch control, with a response latency of ≤20ms; based on edge computing units and spatial computing algorithms, it analyzes user intent in real time and triggers preset interaction rules, and supports custom development interfaces.

[0016] (3) The use of ray tracing and anti-aliasing technology solves the problem of blurred projection edges and ensures image clarity; combined with Micro LED display chip, it achieves HDR effect and 120Hz high refresh rate; based on 5G / Wi-Fi 7 protocol, it supports multi-device image synchronization and scene sharing, such as multi-person collaborative editing of virtual models or online games, with a latency of ≤10ms.

[0017] (4) Supports expansion into scenarios such as games, education, and social interaction. For example, medical students can learn through floating 3D anatomical models, or designers can remotely collaborate to modify virtual architectural drawings. No dedicated server is required. Localized processing is achieved through edge computing and distributed architecture, making it suitable for open environments such as homes, shopping malls, and exhibitions. Detailed Implementation

[0018] Example 1: Immersive Medical Surgery Simulation System

[0019] This embodiment constructs a surgical simulation training platform based on a smart screen metaverse. The system generates high-precision 3D models of human organs in mid-air through a projection module. The model resolution reaches 4K UHD (3840×2160 pixels), supports HDR10+ dynamic lighting rendering, and fine structures such as blood vessels and nerves are clearly visible. The interaction module integrates multimodal sensors. The depth sensor uses ToF technology to achieve 0.8 mm-level operational accuracy (at a distance of 1 meter), which can capture the minute tremors of the doctor's scalpel hand gestures. The microphone array supports multilingual voice command recognition, and doctors can adjust the viewing angle in real time through natural language commands such as "zoom in on blood vessels" and "switch incision angle." The computing module is based on the NVIDIA Jetson Xavier NX chip, completing 5 million soft tissue deformation simulations per second. The physics engine can realistically reproduce blood flow resistance and muscle elasticity. The network module supports 256 medical students simultaneously accessing the same virtual operating room, with a multi-device position synchronization error of less than 3 cm and a latency of less than 8 milliseconds. Experimental data shows that doctors trained using this system have a 40% improvement in operational accuracy and a 25% faster emergency response speed in real surgeries.

[0020] Example 2: Smart City Planning and Disaster Simulation System

[0021] This embodiment is applied to urban planning departments. A projection module projects a 3D city model onto a 5-meter diameter surface in the air. The model includes 100,000 buildings, 200 kilometers of roads, and 5 million square meters of green space, with a ground texture resolution of 0.5 mm / pixel. The interaction module integrates LiDAR and pressure sensors, allowing users to draw new roads or adjust green belts in the air using a handheld laser pointer. The system calculates the rationality of the spatial layout in real time and generates a construction cost estimate report. The computing module adopts a distributed edge computing architecture, with a built-in AI chip boasting 21 TOPS of computing power. It can simultaneously simulate 100 meteorological disaster scenarios (such as torrential rain and flooding, and seismic wave propagation), each scenario containing 500,000 dynamic objects (vehicles, pedestrians, and building components). The network module is based on a 5G millimeter-wave network, supporting simultaneous access for 512 terminal devices. The scene update frequency reaches 120Hz, and the data synchronization error is controlled within 2 cm during multi-department collaboration. In practical applications, this system assisted a city in optimizing its flood prevention plan, improving the accuracy of urban flooding warnings to 92%.

[0022] Example 3: Space Exploration Simulation Training Module

[0023] This embodiment is designed for astronaut training. The projection module generates a 4-meter diameter spherical image of the cabin interior, employing active 3D stereoscopic projection technology with a resolution of 8K (7680×4320 pixels) and supporting a 120Hz high refresh rate. It can simulate the distortion effects of object motion in the microgravity environment of space. The interaction module integrates a brain-computer interface and electromyography (EMG) sensors, allowing astronauts to control the virtual dashboard with their thoughts. The gesture recognition accuracy reaches 0.02 millimeters, simulating the mechanical feedback of grasping floating objects in a zero-gravity environment. The computing module is equipped with a customized AI chip, and the physics engine supports rigid-body-fluid coupling simulation, enabling real-time calculation of the stress distribution of space debris impacting the spacecraft. The network module is based on Starlink satellite communication, achieving two-way data synchronization between Earth and the space station, with a single command transmission latency of less than 50ms. Test data shows that astronauts trained using this system reduced operational errors by 65% ​​during space station docking missions and improved emergency fault handling efficiency by 3 times.

[0024] Example 4: Immersive Film Production and Virtual Shooting

[0025] This embodiment designs a dynamic virtual film set system for the film and television industry. The projection module generates a 270-degree circular stereoscopic image wall with a resolution of 8K×3 (horizontal three-screen splicing), a color gamut covering 140% of the DCI-P3 standard, and a dynamic contrast ratio of 1,500,000:1. The interaction module integrates inertial sensors and optical tracking cameras. Actors wearing motion capture suits can trigger virtual special effects (such as fire and explosions) in real time. Gesture recognition speed reaches 200fps, and voice commands support real-time scheduling by multi-language dubbing directors. The computing module adopts a GPU cluster architecture, rendering 1 billion pixels per second, and supports real-time synthesis of interactive shots between virtual characters and live actors. The network module is based on a private 5G slice network, ensuring that the synchronization error of multi-camera shooting data is less than 1 frame (at 24fps). In actual shooting, this system shortens the post-production special effects cycle by 70% and reduces the cost of virtual scene construction by 85%.

[0026] Example 5: High-end Luxury Customization Experience System

[0027] This embodiment is applied to high-end shopping malls and online stores. The projection module generates a 1.5-meter diameter suspended product image with a 4K UHD resolution, supporting dynamic material switching (such as leather, metal, and fabric) and a texture resolution of 10 micrometers / pixel. The interaction module integrates iris recognition and skin sensors, allowing users to select product colors by gazing and zoom in / out to observe details. The system generates 3D printing preview models in real time. The computing module incorporates an AI style transfer algorithm that automatically generates customized design schemes matching the user's skin tone and body shape, while the physics engine simulates fabric wrinkles and wearing effects. The network module is based on the Wi-Fi 7 protocol, supporting multi-user cross-regional collaborative design with scene synchronization latency of less than 5ms. Data shows that luxury stores using this system saw a 200% increase in average transaction value, a 45% decrease in online return rates, and a 98% user satisfaction rate for virtual try-on.

[0028] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A smart screen meta-border that enables natural human-computer interaction, characterized in that: include: The projection module is used to convert the screen content of smart devices into multi-dimensional stereoscopic images and project them into the air. The interaction module integrates multimodal sensors and motion capture units to collect user behavior data in real time and generate interaction commands; The calculation module runs the dynamic rendering engine and adjusts the physical properties and logical rules of the virtual scene according to instructions. The network module supports data synchronization and scene sharing among multiple devices.

2. The intelligent screen meta-boundary capable of achieving natural human-computer interaction according to claim 1, characterized in that: The projection module includes: The ultra-high-definition display chip has a resolution of 3840×2160 pixels (4K UHD), supports HDR10+ / HLG high dynamic range display, has a brightness of ≥1500 nits, and a contrast ratio of 1,000,000:

1. The optical refraction component uses a 3-layer Fresnel lens array to focus and stereoscopically image aerial images. The focal length can be adjusted from 20 to 100 cm, and the stereoscopic image size is φ200 mm to φ800 mm (supports multiple adjustment levels).

3. The intelligent screen meta-boundary capable of achieving natural human-computer interaction according to claim 1, characterized in that: The interaction module further includes: The depth sensor uses Time-of-Flight (ToF) technology, with a measurement accuracy of ±1mm@1m, a frame rate of 120fps, an effective recognition distance of 5 meters, and supports millimeter-level limb motion capture. The microphone array includes 8 MEMS microphones with a beamforming angle of ±45°, an A-weighted signal-to-noise ratio of 65dB, and supports far-field speech recognition up to 5 meters (including noise suppression and echo cancellation).

4. The intelligent screen meta-boundary capable of achieving natural human-computer interaction according to claim 1, characterized in that: The computing module adopts an edge computing architecture and has a built-in NVIDIA Jetson Xavier NX AI acceleration chip (21 TOPS computing power, 15W power consumption), which supports real-time physics simulation and collision detection (processing 5 million rigid body collisions per second).

5. The intelligent screen meta-boundary capable of achieving natural human-computer interaction according to claim 1, characterized in that: The network module is based on the 5G NR communication protocol (millimeter wave band), with an air interface latency of ≤8ms, a throughput of 10Gbps, and supports synchronization of 256 node devices (scene update frequency of 60Hz).

6. The human-computer interaction method for a smart screen metaverse that enables natural human-computer interaction according to claim 1, characterized in that: Includes the following steps: (1) Step S1: Generate a stereoscopic image of the virtual scene through the projection module (dynamic compensation algorithm eliminates jagged edges and improves image clarity by 40%). (2) Step S2: Use the interaction module to collect user behavior data (gesture trajectory accuracy 0.1mm, voice command recognition accuracy 98%, ambient light intensity monitoring range 10-100,000 lux); (3) Step S3: The calculation module parses the data and triggers the preset interaction logic (AI inference delay ≤ 15ms); (4) Step S4: The network module synchronizes the interaction results to other related devices (multi-device location synchronization error < 5cm).

7. The intelligent screen meta-boundary capable of achieving natural human-computer interaction according to claim 6, characterized in that: In step S2, the user behavior data includes: Gesture trajectory: Supports 26 degrees of freedom gesture recognition, with a recognition speed of within 500ms; Voice commands: Supports Natural Language Understanding (NLU) and multilingual command sets; Ambient light intensity: Dynamically adjust projection brightness (adjustment range 10-100%).

8. The intelligent screen meta-boundary capable of achieving natural human-computer interaction according to claim 1, characterized in that: The ultra-high-definition image processing employs a dynamic compensation algorithm (based on temporal interpolation and motion estimation) and multi-sampling anti-aliasing technology (SSAA+FXAA hybrid mode), which improves image edge sharpness by 60%.

9. The intelligent screen meta-boundary capable of achieving natural human-computer interaction according to claim 1, characterized in that: The virtual scene supports custom development and provides Unity / Unreal engine plugins. Users can define interaction rules and physical rules through a RESTful API (supporting Python / JavaScript script extensions).

10. The intelligent screen meta-boundary capable of achieving natural human-computer interaction according to claim 1, characterized in that: The system is compatible with Android 10+, iOS 14+, and Linux Ubuntu 20.04+ operating systems. It achieves hardware interface unification through the HIDL / DBus driver adaptation layer and supports hot-swappable peripherals (such as external sensors and controllers).