Multimodal human-computer interaction platform and electronic device

WO2026179868A1PCT designated stage Publication Date: 2026-09-03MO XIAODONG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2026/079839
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-25
Filing Date
2026-02-25
Publication Date
2026-09-03

Smart Images

  • Figure CN2026079839_03092026_PF_FP_ABST
    Figure CN2026079839_03092026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a multimodal human-computer interaction platform and an electronic device, comprising: a physical keyboard, a touch and pen-input interactive screen, a stylus, a microphone array, a SLAM sensor, a chip, an operating system, and a docking station integrated into a single human-computer interaction platform; keyboards of three frequently used devices, i.e., a mobile phone, a computer, and a scientific calculator, are uniformly mapped to the physical keyboard; voice interaction, pen interaction, keyboard interaction, touch interaction, and gesture interaction are integrated into a single human-computer interaction platform, and combined with AR glasses and SLAM technology, so as to enable multi-modal interaction in a human-computer interaction environment combining physical and virtual spaces.
Need to check novelty before this filing date? Find Prior Art

Description

A multimodal human-computer interaction platform and electronic device Technical Field

[0001] This invention relates to the fields of spatial computing and human-computer interaction technology, and more particularly to a multimodal human-computer interaction platform and electronic device. Background Technology

[0002] As virtual reality (VR), mixed reality (MR), and augmented reality (AR) technologies extend people's work, study, and entertainment from real space to virtual space, the way humans interact with computers is also changing. VR, MR, and AR are collectively referred to as XR (Extended Reality). XR requires the use of controllers and gestures (XR bare-hand tracking + motion capture) to achieve human-computer interaction. Apple's Vision Pro, released in 2024, ushered in the era of spatial computing. Vision Pro is an all-in-one MR (Mixed Reality) headset that uses VST (Video See Through) technology to achieve virtual-real fusion. VST (Video See Through) superimposes virtual images onto a real-time video stream in the real world, allowing users to view augmented reality content through a head-mounted display, thus achieving virtual-real fusion. Unlike MR headsets, AR (Augmented Reality) glasses use OST technology to achieve augmented reality. OST (Optical See Through) projects virtual images into the user's field of vision through a transparent optical display module, allowing the user to see both the real world and augmented reality content simultaneously. At the experiential level, the perceptual differences between these two technologies are very significant. With VST, the user ultimately sees a video combining real and virtual elements, while with OST, the user sees a real world mixed with virtual content. The Vision Pro head-mounted MR all-in-one device integrates up to 23 sensors, a high-performance computing chip, and the Vision OS operating system. These include a microphone array, inertial measurement unit (IMU), optical sensors, infrared sensors, and LiDAR. It relies on these sensors to achieve VST and uses the data collected by the sensors to perform SLAM (Simultaneous Localization and Mapping) and spatial computing. SLAM is a technology that simultaneously achieves device localization and environmental mapping. The principle of SLAM is to use sensors such as depth cameras, LiDAR, and inertial measurement units (IMUs) to collect environmental information, and then use algorithms to fuse this information to determine the device's position in an unknown environment and build an environmental map. SLAM technology is divided into LiDAR SLAM and visual SLAM. Lidar SLAM uses 2D or 3D lidar, while visual SLAM uses optical and infrared sensors, currently mainly including monocular cameras, binocular cameras, and RGBD cameras. SLAM technology gives Vision Pro a very good gesture interaction experience, but because Vision Pro integrates all components into a single head-mounted device, its overall weight is over 600 grams, making it very uncomfortable to wear for extended periods.AR glasses adopt a split-type technical approach, retaining only the optical engine (optical projection imaging unit), inertial measurement unit (IMU), and a small number of video and audio components. Separating the sensors, chips, and operating system from the head-mounted device significantly reduces weight, generally not exceeding 100 grams, allowing for extended wear. However, this also brings another problem: the lack of sensors such as depth cameras to collect environmental information makes it difficult to achieve gesture interaction based on SLAM technology. XR controllers play an important role in MR and AR interaction. While XR controllers or gestures (XR hand tracking + motion capture) enable various interactions in virtual space, inputting text, numbers, and punctuation marks using an XR controller (appearing as a beam in virtual space) on a virtual keyboard (phone keyboard / computer keyboard / calculator keyboard) or using gestures (XR hand tracking + motion capture) results in a significantly less tactile experience compared to using a mouse on a PC screen. Unlike mobile touchscreen keyboards, XR keyboards lack haptic feedback, making the user experience worse and less efficient. It's better to map the three most frequently used keyboards—phone, computer, and scientific calculator—to physical keyboards instead of using XR virtual keyboards. The haptic feedback of a physical keyboard and the ability to write and draw on a flat interactive screen with a stylus are irreplaceable by other XR interaction methods. For MR headsets and AR glasses to become productivity tools, they rely on a combination of physical keyboards, styluses, and interactive screens. Today's smartphones, tablets, and laptops are becoming increasingly thinner and lighter, but this leads to another problem: a general lack of expansion interfaces. The technical solution presented in this invention integrates voice interaction, pen interaction, keyboard interaction, touch interaction, and gesture interaction into a single human-computer interaction platform, achieving multimodal interaction. It also integrates multiple sensors, chips, and expansion interfaces, relying on SLAM technology and spatial computing to achieve gesture interaction in conjunction with AR glasses. In summary, the technical solution presented in this invention effectively solves the problems encountered by MR headsets and AR glasses in human-computer interaction. Technical issues

[0003] As virtual reality (VR), mixed reality (MR), and augmented reality (AR) technologies extend people's work, study, and entertainment from real space to virtual space, the way humans interact with computers is also changing. VR, MR, and AR are collectively referred to as XR (Extended Reality). XR requires the use of controllers and gestures (XR bare-hand tracking + motion capture) to achieve human-computer interaction. Apple's Vision Pro, released in 2024, ushered in the era of spatial computing. Vision Pro is an all-in-one MR (Mixed Reality) headset that uses VST (Video See Through) technology to achieve virtual-real fusion. VST (Video See Through) superimposes virtual images onto a real-time video stream in the real world, allowing users to view augmented reality content through a head-mounted display, thus achieving virtual-real fusion. Unlike MR headsets, AR (Augmented Reality) glasses use OST technology to achieve augmented reality. OST (Optical See Through) projects virtual images into the user's field of vision through a transparent optical display module, allowing the user to see both the real world and augmented reality content simultaneously. At the experiential level, the perceptual differences between these two technologies are very significant. With VST, the user ultimately sees a video combining real and virtual elements, while with OST, the user sees a real world mixed with virtual content. The Vision Pro head-mounted MR all-in-one device integrates up to 23 sensors, a high-performance computing chip, and the Vision OS operating system. These include a microphone array, inertial measurement unit (IMU), optical sensors, infrared sensors, and LiDAR. It relies on these sensors to achieve VST and uses the data collected by the sensors to perform SLAM (Simultaneous Localization and Mapping) and spatial computing. SLAM is a technology that simultaneously achieves device localization and environmental mapping. The principle of SLAM is to use sensors such as depth cameras, LiDAR, and inertial measurement units (IMUs) to collect environmental information, and then use algorithms to fuse this information to determine the device's position in an unknown environment and build an environmental map. SLAM technology is divided into LiDAR SLAM and visual SLAM. Lidar SLAM uses 2D or 3D lidar, while visual SLAM uses optical and infrared sensors, currently mainly including monocular cameras, binocular cameras, and RGBD cameras. SLAM technology gives Vision Pro a very good gesture interaction experience, but because Vision Pro integrates all components into a single head-mounted device, its overall weight is over 600 grams, making it very uncomfortable to wear for extended periods.AR glasses adopt a split-type technical approach, retaining only the optical engine (optical projection imaging unit), inertial measurement unit (IMU), and a small number of video and audio components. Separating the sensors, chips, and operating system from the head-mounted device significantly reduces weight, generally not exceeding 100 grams, allowing for extended wear. However, this also brings another problem: the lack of sensors such as depth cameras to collect environmental information makes it difficult to achieve gesture interaction based on SLAM technology. XR controllers play an important role in MR and AR interaction. While XR controllers or gestures (XR hand tracking + motion capture) enable various interactions in virtual space, inputting text, numbers, and punctuation marks using an XR controller (appearing as a beam in virtual space) on a virtual keyboard (phone keyboard / computer keyboard / calculator keyboard) or using gestures (XR hand tracking + motion capture) results in a significantly less tactile experience compared to using a mouse on a PC screen. Unlike mobile touchscreen keyboards, XR keyboards lack haptic feedback, making the user experience worse and less efficient. It's better to map the three most frequently used keyboards—phone, computer, and scientific calculator—to physical keyboards instead of using XR virtual keyboards. The haptic feedback of a physical keyboard and the ability to write and draw on a flat interactive screen with a stylus are irreplaceable by other XR interaction methods. For MR headsets and AR glasses to become productivity tools, they rely on a combination of physical keyboards, styluses, and interactive screens. Today's smartphones, tablets, and laptops are becoming increasingly thinner and lighter, but this leads to another problem: a general lack of expansion interfaces. The technical solution presented in this invention integrates voice interaction, pen interaction, keyboard interaction, touch interaction, and gesture interaction into a single human-computer interaction platform, achieving multimodal interaction. It also integrates multiple sensors, chips, and expansion interfaces, relying on SLAM technology and spatial computing to achieve gesture interaction in conjunction with AR glasses. In summary, the technical solution presented in this invention effectively solves the problems encountered by MR headsets and AR glasses in human-computer interaction. Technical solutions

[0004] The purpose of this invention is to provide a multimodal human-computer interaction technical solution to solve the problems encountered in XR human-computer interaction. It unifies the mapping of three frequently used keyboards—mobile phones, computers, and scientific calculators—to a physical keyboard, and integrates voice interaction, pen interaction, keyboard interaction, touch interaction, and gesture interaction (XR bare-hand tracking + motion capture) into a single human-computer interaction platform to achieve multimodal interaction. It also integrates multiple SLAM sensors, chips, and operating systems, as well as expansion interfaces (docking stations), and relies on SLAM technology and spatial computing to achieve gesture interaction in conjunction with AR glasses.

[0005] The technical solution adopted in this invention includes the following:

[0006] This platform integrates a physical keyboard, a handwriting-enabled touchscreen, a stylus, a microphone array, a SLAM sensor, a chip, an operating system, and a docking station into a single human-computer interaction platform. It maps three frequently used keyboards—mobile phones, computers, and scientific calculators—to a physical keyboard, and integrates voice interaction, pen interaction, keyboard interaction, touch interaction, and gesture interaction (XR bare-hand tracking + motion capture) into a single human-computer interaction platform. When used with AR glasses, it enables multimodal interaction in a human-computer interaction environment that combines real and virtual spaces.

[0007] On the other hand, the present invention also provides an electronic device, comprising:

[0008] Main unit;

[0009] A processor is disposed in the main unit;

[0010] The aforementioned multimodal human-computer interaction platform is connected to the processor. Beneficial effects

[0011] Compared with existing technologies, the beneficial effects of this invention are: it unifies the mapping of three frequently used keyboards—mobile phones, computers, and scientific calculators—to a physical keyboard. It integrates voice interaction, pen interaction, keyboard interaction, touch interaction, and gesture interaction (XR bare-hand tracking + motion capture) into a single human-computer interaction platform, achieving multimodal interaction. An integrated docking station facilitates user connection to various devices, providing a better experience. The multimodal human-computer interaction platform combined with AR glasses combines real and virtual spaces, enabling large-screen + multi-screen + gesture interaction, one-click control of multiple screens, and one-screen control of multiple screens, significantly improving interaction efficiency. Attached Figure Description

[0012] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, the following describes the preferred embodiments of the present invention in detail with reference to the accompanying drawings.

[0013] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0014] Figure 1 shows the first embodiment of the present invention.

[0015] Figure 2 shows the second embodiment of the present invention.

[0016] Figure 3 shows the third embodiment of the present invention.

[0017] Figure 4 shows a top view of the third embodiment.

[0018] Figure 5 shows a bottom view of three embodiments of the present invention.

[0019] Figure 6 shows a schematic diagram illustrating the working principle of the present invention. The best embodiment of the present invention

[0020] Figure 2 shows a schematic diagram of the second embodiment of the present invention, including: a handwriting touch screen 1, a docking station 2, a microphone 3, a stylus 4, a single-lens optical camera 5, a microphone 6 (which, together with microphone 3, forms a microphone array), a single-lens optical camera 7 (which, together with single-lens optical camera 5, forms a dual-lens optical camera), and a single-key physical keyboard 8. The planar interactive screen uses a 12-inch handwriting touch screen, which interacts collaboratively with multiple vertical screens in real and virtual spaces, allowing one screen to control multiple screens. The microphone array has far-field sound pickup and noise reduction capabilities. The stylus uses an active capacitive pen or an electromagnetic pen, with a lanyard attached to the end to prevent loss; it is placed in a slot when not in use. The binocular camera works in conjunction with the SLAM algorithm to achieve gesture interaction. SLAM is a technology for Simultaneous Localization and Mapping. The binocular optical camera acts as a visual SLAM sensor. Binocular SLAM uses the parallax of the left and right eyes to calculate the distance between pixels, thereby achieving its own localization. Stereo vision can estimate depth both in motion and at rest, eliminating the problem of monocular vision being unable to obtain depth information. The camera is mounted on a multimodal human-computer interaction platform and can rotate left and right to facilitate selecting a suitable angle for gesture interaction. The multimodal human-computer interaction platform has its own CPU and operating system, integrating keyboard interaction, pen interaction, voice interaction, touchscreen interaction, and gesture interaction (XR bare-hand tracking + motion capture) to achieve multimodal interaction in conjunction with XR. AR glasses differ from MR headsets; to achieve lightweight design, SLAM sensors, chips, and operating systems need to be integrated into the multimodal human-computer interaction platform. Embodiments of the present invention

[0021] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0022] Figure 1 shows a schematic diagram of the first embodiment of the present invention, including: a handwriting touch screen 1, a docking station 2, a microphone 3, a stylus 4, a monocular optical camera 5, a microphone 6 (which, together with microphone 3, forms a microphone array), and a single-key physical keyboard 7. The planar interactive screen uses a 12-inch handwriting touch screen, which interacts collaboratively with multiple vertical screens in real and virtual spaces, allowing one screen to control multiple screens. The microphone array has far-field sound pickup and noise reduction capabilities. The stylus is an active capacitive pen or an electromagnetic pen, with a lanyard attached to the end to prevent loss; it is placed in a slot when not in use. The monocular camera works in conjunction with the SLAM algorithm to achieve gesture interaction. SLAM is a technology for simultaneous localization and mapping (SLM). The monocular optical camera, as a visual SLAM sensor, has the advantages of being inexpensive and lightweight, but its disadvantage is that it cannot obtain depth information. The monocular camera is fixed to the multimodal human-computer interaction platform by a bracket and can rotate left and right to facilitate the selection of a suitable angle for gesture interaction. The multimodal human-computer interaction platform integrates its own CPU and operating system, unifying keyboard, pen, voice, touchscreen, and gesture interaction (XR bare-hand tracking + motion capture) to achieve multimodal interaction in conjunction with XR. Unlike MR headsets, AR glasses require lightweight design, necessitating the integration of SLAM sensors, chips, and the operating system into the multimodal human-computer interaction platform. This platform maps frequently used keyboards (phone, computer, and scientific calculator) to a single physical keyboard, enabling one-click control of multiple screens and one-screen control of multiple screens, significantly improving the efficiency of human-computer interaction in both real and virtual spaces.

[0023] Figure 2 shows a schematic diagram of the second embodiment of the present invention, including: a handwriting touch screen 1, a docking station 2, a microphone 3, a stylus 4, a single-lens optical camera 5, a microphone 6 (which, together with microphone 3, forms a microphone array), a single-lens optical camera 7 (which, together with single-lens optical camera 5, forms a dual-lens optical camera), and a single-key physical keyboard 8. The planar interactive screen uses a 12-inch handwriting touch screen, which interacts collaboratively with multiple vertical screens in real and virtual spaces, allowing one screen to control multiple screens. The microphone array has far-field sound pickup and noise reduction capabilities. The stylus uses an active capacitive pen or an electromagnetic pen, with a lanyard attached to the end to prevent loss; it is placed in a slot when not in use. The binocular camera works in conjunction with the SLAM algorithm to achieve gesture interaction. SLAM is a technology for Simultaneous Localization and Mapping. The binocular optical camera acts as a visual SLAM sensor. Binocular SLAM uses the parallax of the left and right eyes to calculate the distance between pixels, thereby achieving its own localization. Stereo vision can estimate depth both in motion and at rest, eliminating the problem of monocular vision being unable to obtain depth information. The camera is mounted on a multimodal human-computer interaction platform and can rotate left and right to facilitate selecting a suitable angle for gesture interaction. The multimodal human-computer interaction platform has its own CPU and operating system, integrating keyboard interaction, pen interaction, voice interaction, touchscreen interaction, and gesture interaction (XR bare-hand tracking + motion capture) to achieve multimodal interaction in conjunction with XR. AR glasses differ from MR headsets; to achieve lightweight design, SLAM sensors, chips, and operating systems need to be integrated into the multimodal human-computer interaction platform.

[0024] Figure 3 shows a schematic diagram of the third embodiment of the present invention, including: a handwriting touch screen 1, a docking station 2, a microphone 3, a stylus 4, an optical plus infrared depth camera 5 (RGBD depth camera), a microphone 6 (which, together with microphone 3, forms a microphone array), and a single-key physical keyboard 7. The planar interactive screen uses a 12-inch handwriting touch screen, which interacts collaboratively with multiple vertical screens in real and virtual spaces, allowing one screen to control multiple screens. The microphone array has far-field sound pickup and noise reduction capabilities. The stylus uses an active capacitive pen or an electromagnetic pen, with a lanyard attached to the end to prevent loss; it is placed in a slot when not in use. The depth camera shown in the figure is in a horizontal position. The depth camera's bracket can be adjusted to a 90-degree tilt angle. When used with a PC, to avoid obstructing the PC monitor and affecting the display effect, the depth camera's bracket can be adjusted to a horizontal position. The RGBD camera uses an infrared camera to acquire three-dimensional spatial depth information. Its characteristic is that it directly measures the distance of each pixel in the image from the camera through infrared structured light or Time-of-Flight principles, achieving more accurate synchronous positioning and spatial calculation.

[0025] Figure 4 shows a top view of the third embodiment, including: a handwriting touch screen 1, a docking station 2, a microphone 3, a stylus 4, an optical plus infrared depth camera 5 (RGBD camera), a microphone 6 (forming a microphone array together with microphone 3), and a single-key physical keyboard 7. The microphone array has far-field sound pickup and noise reduction capabilities. The depth camera shown in the figure is in a vertical standing position (the position state of the depth camera when working). When used with AR glasses to achieve gesture interaction, the depth camera's stand is adjusted to a vertical standing position. The depth camera located at the top of the stand can rotate left and right to facilitate selecting a suitable angle for gesture interaction. The XR virtual keyboard is different from the virtual keyboard on a mobile phone touchscreen; it lacks haptic feedback, resulting in a poor user experience and lower input efficiency than a physical keyboard. The multimodal human-computer interaction platform maps the three frequently used keyboards—mobile phone, computer, and scientific calculator—to a physical keyboard, enabling one-click control of multiple screens and one screen control of multiple screens, significantly improving the efficiency of human-computer interaction combining real and virtual spaces.

[0026] Figure 5 shows bottom views of the first, second, and third embodiments of the present invention, with the specific location of the docking station 2 marked by dashed lines. Corresponding to the docking station 2 in the above three embodiments (as seen in the rear view), since the docking station 2 is located in the rear view position of the multimodal human-computer interaction platform, it is not visible from the perspective of the first, second, and third embodiments, nor from the top view of the third embodiment in Figure 4. Combining the bottom and rear views (partial), it can be seen that the docking station contains multiple USB ports. As mobile phones, tablets, and laptops become increasingly thinner and more portable, the lack of expansion ports has become a pain point for users. The multimodal human-computer interaction platform integrates a docking station, facilitating users to connect various devices and bringing a better experience.

[0027] Figure 6 is a schematic diagram illustrating the working principle of three embodiments of the present invention. As shown in Figure 6, it is represented in the form of a general-purpose computing device. Its components include: an operating system 1, various application programs 2, ROM (Read-Only Memory) 3, a storage unit 4, RAM (Random Access Memory) 5, a cache 6, a processor unit 7, a bus connecting different components 8, an I / O interface 9, a display unit 10, and a network adapter 11. The storage unit stores the operating system and application program code, which is executed by the processing unit 7. The processing unit 7 can execute various applications including SLAM algorithms to achieve gesture interaction (XR bare-hand tracking + motion capture) and multimodal interaction. The keyboard, pen, touchscreen, sound sensor (microphone array), and visual SLAM sensor (optical camera and infrared camera) are all external devices. The multimodal human-computer interaction platform has its own CPU and operating system. Although it can be used independently as a terminal device (equivalent to a PC), its main task is to work with AR glasses to achieve multimodal human-computer interaction. AR glasses differ from MR headsets; to achieve lightweight design, the SLAM sensor, chip, and operating system need to be integrated into the multimodal human-computer interaction platform.

[0028] The three embodiments provided by this invention all have a complex form with a real part and an imaginary part. The real part corresponds to the physical keyboard, and the imaginary part corresponds to the planar interactive screen. Combined with AR glasses and SLAM technology, multimodal interaction is achieved in human-computer interaction that combines real and virtual spaces, enabling one-click control of multiple screens and one screen to control multiple screens simultaneously, significantly improving interaction efficiency. SLAM technology is divided into LiDAR SLAM and visual SLAM. LiDAR SLAM uses 2D or 3D LiDAR, while visual SLAM uses optical and infrared sensors, mainly monocular cameras, binocular cameras, and RGBD cameras. The above three embodiments use the inexpensive and lightweight visual SLAM, omitting LiDAR SLAM, but this does not mean that LiDAR SLAM is not used in all embodiments. The above embodiments all have the following technical effects: unifying the mapping of three frequently used keyboards—mobile phones, computers, and scientific calculators—to a physical keyboard; integrating voice interaction, pen interaction, keyboard interaction, touch interaction, and gesture interaction (XR bare-hand tracking + motion capture) into a single human-computer interaction platform to achieve multimodal interaction. An integrated docking station allows users to easily connect various devices for a better experience. The multimodal human-computer interaction platform combined with AR glasses, leveraging SLAM for gesture interaction, integrates real and virtual spaces, enabling large-screen + multi-screen + gesture interaction. One-click control of multiple screens and one-screen control of multiple screens significantly improves interaction efficiency.

[0029] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention. Industrial applicability

[0030] This invention has strong industrial applicability. It maps three frequently used keyboards—phones, computers, and scientific calculators—to a unified physical keyboard; it integrates voice interaction, pen interaction, keyboard interaction, touch interaction, and gesture interaction (XR bare-hand tracking + motion capture) into a single human-computer interaction platform, achieving multimodal interaction. An integrated docking station facilitates user connection to multiple devices, providing a better experience. The multimodal human-computer interaction platform combined with AR glasses combines real and virtual spaces, enabling large-screen + multi-screen + gesture interaction, one-click control of multiple screens, and one screen control of multiple screens, significantly improving interaction efficiency. Instead of using an XR virtual keyboard, it maps the three frequently used keyboards—phones, computers, and scientific calculators—to a unified physical keyboard. The tactile feedback of the physical keyboard and the handwriting / drawing on the planar interactive screen in real space are irreplaceable by other XR interaction methods. For MR headsets and AR glasses to become productivity tools, they cannot function without a physical keyboard + stylus + planar interactive screen. Today's phones, tablets, and laptops are becoming increasingly thin and portable, but this leads to another problem: a general lack of expansion interfaces. The technical solution provided by this invention integrates voice interaction, pen interaction, keyboard interaction, touch interaction, and gesture interaction into a single human-computer interaction platform, achieving multimodal interaction. It also integrates multiple sensors, chips, and expansion interfaces, relying on SLAM technology and spatial computing to achieve gesture interaction in conjunction with AR glasses. In summary, the technical solution provided by this invention effectively solves the problems encountered by MR headsets and AR glasses in human-computer interaction. Sequence List Free Content

[0031] Type the free content description paragraph for the sequence list here.

Claims

1. A multimodal human-computer interaction platform, characterized in that, include: This platform integrates a physical keyboard, a handwriting-enabled touchscreen, a stylus, a microphone array, a SLAM sensor, a chip, an operating system, and a docking station into a single human-computer interaction platform. It also maps three frequently used keyboards—mobile phones, computers, and scientific calculators—to a physical keyboard. Furthermore, it integrates voice interaction, pen interaction, keyboard interaction, touch interaction, and gesture interaction into a single platform. By combining AR glasses and SLAM technology, it enables multimodal interaction in a human-computer interaction environment that blends real and virtual spaces.

2. The multimodal human-computer interaction platform according to claim 1, characterized in that, This multimodal human-computer interaction platform uses a single-lens optical camera, combined with AR glasses and SLAM technology, to achieve gesture interaction.

3. The multimodal human-computer interaction platform according to claim 1, characterized in that, This multimodal human-computer interaction platform uses binocular optical cameras, combined with AR glasses and SLAM technology, to achieve gesture interaction.

4. The multimodal human-computer interaction platform according to claim 1, characterized in that, This multimodal human-computer interaction platform uses an RGBD depth camera, combined with AR glasses and SLAM technology, to achieve gesture interaction.

5. The multimodal human-computer interaction platform according to claim 1, characterized in that, This multimodal human-computer interaction platform integrates a docking station, making it convenient for users to connect various devices.

6. The multimodal human-computer interaction platform according to claim 1, characterized in that, This multimodal human-computer interaction platform comes with its own CPU and operating system. Combined with AR glasses and SLAM technology, it enables multimodal interaction in a human-computer interaction environment that combines real and virtual spaces.

7. The multimodal human-computer interaction platform according to claim 4, characterized in that, The RGBD depth camera on this multimodal human-computer interaction platform uses an adjustable pitch angle bracket. The RGBD depth camera is fixed on the top of the bracket and can rotate left and right to facilitate the selection of a suitable angle for gesture interaction.

8. An electronic device, characterized in that, include: Main unit; A processor is disposed in the main unit; A multimodal human-computer interaction platform as described in any one of claims 1-7, connected to the processor.