An MR headset intelligent control system
By using high-definition cameras and depth sensors in the MR headset to automatically adjust parameters, combined with voice recognition and eye movement filtering technology, the misjudgment and fatigue problems of gesture, voice and eye movement interaction control in complex environments of the MR headset is solved, achieving higher interaction accuracy and user experience fluency.
Patent Information
- Application Number
- CN202411566016.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2044-11-05
AI Technical Summary
The existing MR headsets have misjudgment and fatigue problems in gesture, voice and eye movement interaction control in complex environments, affecting the user experience.
The high-definition camera and depth sensor are used to automatically adjust parameters, combined with voice recognition and eye movement filtering technology, and accurate interactive control is achieved through a collaborative linkage module.
Improves the accuracy of gesture, voice and eye movement interactions, reduces user fatigue, and enhances immersion and smooth operation.
Smart Images

Figure CN119512371B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of MR headsets, and more specifically, to an intelligent control system for an MR headset. Background Art
[0002] An MR (Mixed Reality) headset is an advanced head-mounted display device that combines the characteristics of virtual reality (VR) and augmented reality (AR), enabling users to simultaneously perceive information from the real world and the virtual world and interact with them.
[0003] MR headsets are usually equipped with a variety of sensors for perceiving the position and motion state of the user's head, as well as the three-dimensional spatial information of the surrounding environment. For example, the IMU can be used to obtain the rotation and translation data of the head, and cameras and depth sensors can scan the surrounding environment to construct a spatial model of the environment, thereby determining the position and presentation method of virtual objects in the real space. At the same time, it uses special display technology to project virtual images into the user's field of vision while ensuring that the user can see the real world. Some MR headsets achieve this function through semi-transparent optical lenses, allowing the light of the real world to pass through the lenses, while virtual images are superimposed on the real scene seen by the user through optical reflection or projection. In addition, some use video see-through technology, that is, the camera captures the picture of the real world, and then the virtual content is synthesized with the picture and then displayed to the user.
[0004] Currently, the following are some drawbacks of the interaction control methods of MR headsets:
[0005] 1) Drawbacks of gesture interaction control: In actual use, for some relatively complex or similar gestures, the recognition system of the headset may make misjudgments. For example, when the user quickly makes some hand movements, or the amplitude, angle, etc. of the hand movements deviate from the preset standard gestures to a certain extent, the system may not be able to accurately recognize the user's intention, resulting in operation errors. At the same time, factors such as the light in the surrounding environment and the occlusion of objects may affect the accuracy of gesture recognition. For example, in an environment with too strong or too dim light, the camera may not be able to clearly capture the user's hand movements; if the user's hand is blocked by other objects, the system cannot correctly recognize the gesture.
[0006] 2) Disadvantages of voice interaction control: In a noisy environment, the microphone of the headset may receive a large amount of background noise, which will affect the accuracy of voice recognition. For example, in public places, factory workshops and other environments, surrounding voices, machine sounds and other noises may cause the system to fail to correctly recognize the user's voice commands. At the same time, there are differences in accents and language habits: Users in different regions may have different accents and language habits, which may cause the voice recognition system to inaccurately recognize the voice commands of some users. For example, the pronunciation of some local dialects may be quite different from the standard Mandarin pronunciation, and the system may not be able to correctly recognize these dialect voice commands.
[0007] 3) Disadvantages of eye movement interaction control: When using the eye movement tracking function for a long time, the user's eyes may feel fatigued and uncomfortable. Because the user needs to continuously fixate on specific areas or objects in the headset to interact, this will keep the eye muscles in a tense state, and long-term use may cause problems such as dry and painful eyes. At the same time, the eye movement tracking technology may be affected by the user's involuntary eye movements, blinks and other actions, resulting in false triggers or incorrect operations. For example, when the user is thinking or inadvertently looks in a certain direction, the system may misjudge it as the user's operation instruction and thus perform some unnecessary operations.
[0008] For the problems in the related technologies, no effective solutions have been proposed yet. Summary of the Invention
[0009] For the problems in the related technologies, the present invention proposes an intelligent control system for an MR headset to overcome the above-mentioned technical problems existing in the existing related technologies.
[0010] The technical solution of the present invention is realized as follows:
[0011] An intelligent control system for an MR headset includes: a gesture interaction control module, a voice interaction control module, an eye movement interaction control module and a collaborative linkage module, wherein;
[0012] The gesture interaction control module is used to configure the high-definition camera and the depth sensor to automatically adjust parameters in a complex environment and provide real-time gesture prompts and feedback for the recognition and instruction control of complex gestures; wherein, the high-definition camera is used to collect hand movements; the depth sensor is used to embed a gesture recognition algorithm to adapt to different hand movements, speeds and angles to meet the recognition of complex gestures;
[0013] The voice interaction control module is used to filter and suppress the noise in the environment, highlight the user's voice signal, and automatically adjust and optimize according to the user's voice characteristics through the embedded voice recognition engine to improve the recognition of voice control instructions in a noisy environment. It includes: a noise elimination module and a voice enhancement optimization module, wherein the noise elimination module is used to filter and suppress the noise in the environment by using a microphone array and digital signal processing to highlight the user's voice signal; the voice enhancement optimization module is used to embed the voice recognition engine, and can automatically adjust and optimize according to the user's voice characteristics;
[0014] The eye movement interaction control module is used to track the user's eye movement and perform filtering and command control during the recognition process of blinking actions, including: a calibration tracking module and a blink detection filtering module, wherein the calibration tracking module is used to perform eye movement calibration according to the eye movement tracking algorithm before using the head display; the blink detection filtering module is used to determine whether it is a blinking action by analyzing the frequency and amplitude characteristics of the eye movement, and perform filtering during the recognition process;
[0015] Furthermore, it also includes:
[0016] The collaborative linkage module is used for the gesture interaction control module, the voice interaction control module and the eye movement interaction control module to work together, and when the gesture interaction control module, the voice interaction control module or the eye movement interaction control module performs priority control, the gesture interaction control module, the voice interaction control module or the eye movement interaction control module is linked for collaborative control.
[0017] Furthermore, the gesture interaction control module further includes: an environment calibration template and a gesture guidance module, wherein;
[0018] The environmental calibration template is used to enable the head display to automatically adjust parameters in different usage environments to adapt to environmental changes, which at least includes: automatically adjusting the exposure parameters of the high-definition camera in an environment with too strong or too dark light to ensure that hand movements can be clearly captured;
[0019] The gesture guidance module is used to provide real-time gesture prompts and feedback in the head-mounted display interface to help users perform gesture operations correctly.
[0020] Furthermore, the voice interaction control module further includes: a voice setting module and a voice feedback module, wherein;
[0021] The voice setting module is used to provide voice templates and shortcut command settings, allowing users to customize voice commands;
[0022] The voice feedback module is used to provide voice feedback after recognizing the user's voice command.
[0023] Furthermore, the blink detection and filtering module includes the following steps:
[0024] Previously, the differential algorithm is used to calculate the first-order differential velocity and second-order differential acceleration of the eye movement position data to obtain the velocity and acceleration of the eye movement.
[0025] The Canny edge detection image recognition algorithm is used to determine the boundary of the pupil, calculate the pupil area, and obtain the change in the pupil area.
[0026] The eye movement velocity, acceleration, and change in pupil area are used as feature inputs, and a support vector machine is used to determine whether it is a blink action.
[0027] Calibration is used in eye movement tracking to select a preset scenario. If it is determined to be a blink action, the blink action is filtered and a feedback prompt is given.
[0028] Advantages of the present invention:
[0029] Through the gesture interaction control module, the present invention realizes accurate recognition of complex gestures and interactive recognition of command control, enabling users to interact with the virtual and real integrated environment more naturally and smoothly. At the same time, when the voice interaction control module and the eye movement interaction control module achieve accurate interactive recognition, the user's attention can be more concentrated on the experience content itself rather than being disturbed by frequent recognition errors.
[0030] The accurate interactive recognition of the present invention reduces the fatigue of users caused by repeated operations or correcting incorrect operations. For the recognition of complex gestures and command control, users do not need to repeat the same action multiple times to achieve the purpose, thus reducing the burden on the arm and hand muscles. Similarly, in improving the recognition of voice control commands in a noisy environment, accurate recognition means that users do not need to speak loudly or slowly repeatedly for the system to understand their intentions, thereby reducing the fatigue caused by voice interaction. For eye movement interaction control, during the recognition process of blink actions, filtering and command control are performed. Accurate recognition can prevent users from constantly adjusting their line of sight due to frequent false triggers, reducing eye strain and fatigue. Description of the Drawings
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0032] Figure 1 It is a schematic block diagram of an MR head-mounted display intelligent control system according to an embodiment of the present invention. Detailed implementation manners
[0033] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention belong to the scope of protection of the present invention.
[0034] According to an embodiment of the present invention, an MR headset intelligent control system is provided.
[0035] As Figure 1 shown, the MR headset intelligent control system according to an embodiment of the present invention includes: a gesture interaction control module 1, a voice interaction control module 2, an eye movement interaction control module 3, and a collaborative linkage module 4, wherein;
[0036] The gesture interaction control module 1 is used to configure the high-definition camera 11 and the depth sensor 12 to automatically adjust parameters in a complex environment and provide real-time gesture prompts and feedback, and perform recognition and instruction control of complex gestures, wherein;
[0037] The high-definition camera 11 is used to collect hand movements;
[0038] The depth sensor 12 is used to embed a gesture recognition algorithm so that it can better adapt to different hand movements, speeds and angles, and meet the recognition ability of complex gestures.
[0039] In this technical solution, the high-definition camera 11 and the depth sensor 12 are adopted to more accurately capture the position, shape and movement trajectory of the hand, and improve its accuracy and sensitivity. At the same time, the depth sensor 12 can accurately measure the distance between the object and the headset under different lighting conditions, thereby improving the accuracy of gesture recognition.
[0040] The environment calibration template 13 is used to enable the headset to automatically adjust parameters in different usage environments to adapt to environmental changes.
[0041] Specifically, in this technical solution, in an environment with too strong or too dark light, the exposure parameters of the high-definition camera are automatically adjusted to ensure that hand movements can be clearly captured.
[0042] The gesture guidance module 14 is used to provide real-time gesture prompts and feedback in the interface of the headset to help users perform gesture operations correctly.
[0043] In this technical solution, when the user's gesture is not recognized, the system can display a prompt message to tell the user how to adjust the gesture. Meanwhile, during application, a detailed gesture operation guide and training tutorial can be provided to the user in advance, enabling the user to understand the correct gesture operation methods and precautions. For example, through video tutorials, graphic illustrations, etc., show the user how to make standard gesture actions and how to avoid some common incorrect operations.
[0044] The voice interaction control module 2 is used to filter and suppress the noise in the environment, highlight the user's voice signal, and automatically adjust and optimize according to the user's voice characteristics through the embedded voice recognition engine, improving the recognition of voice control commands in a noisy environment, including: a noise cancellation module 21 and a voice enhancement and optimization module 22, where;
[0045] The noise cancellation module 21 is used to filter and suppress the noise in the environment by using a microphone array and digital signal processing, highlighting the user's voice signal.
[0046] The voice enhancement and optimization module 22 is used to embed a voice recognition engine, which can automatically adjust and optimize according to the user's voice characteristics.
[0047] In this technical solution, through the voice recognition engine, the adaptability to different accents, speech rates, and language habits can be improved, and combined with the noise cancellation module, the voice recognition accuracy in a noisy environment can be increased.
[0048] The voice setting module 23 is used to provide voice templates and shortcut command settings for the user to customize voice commands;
[0049] In this technical solution, through the voice setting module, it is convenient for the user to perform personalized voice settings, enabling the system to better adapt to the user's voice characteristics. For example, the user can record their own voice samples according to the provided voice templates and shortcut command settings for the system to learn and optimize, improving the recognition accuracy of the user's specific voice and the interaction efficiency. For example, the user can set shortcut commands such as "open the picture" and "switch scenes" for quick operations during use.
[0050] The voice feedback module 24 is used to provide voice feedback after recognizing the user's voice command.
[0051] In this technical solution, by introducing the voice feedback module, the user can know whether the system has correctly understood their command. Specifically, for example: after the user issues the command "enlarge the image", the system can reply "enlarging the image" for the user to confirm whether the operation is correct.
[0052] In addition, for some important operations or instructions that may be ambiguous, the user is required to confirm. For example, when the system recognizes that there may be multiple interpretations of the user's voice command, it can ask the user "Do you want to open File A or File B?" and perform the operation after the user confirms.
[0053] The eye movement interaction control module 3 is used to track the user's eye movement and perform filtering and instruction control during the recognition of the blinking action, including: a calibration tracking module 31 and a blinking detection and filtering module 32, where;
[0054] The calibration tracking module 31 is used to perform eye movement calibration according to the eye movement tracking algorithm before using the head-mounted display.
[0055] In this technical solution, through the eye movement interaction control module 3, it is ensured that the system can accurately track the user's eye movement. For example, let the user fix their gaze on a series of specific points or patterns, and the system calibrates according to the user's gaze position to improve the accuracy of eye movement tracking.
[0056] At the same time, by introducing the eye movement tracking algorithm, the accuracy and stability of tracking are improved. The system can automatically learn the user's eye movement habits, so as to better predict the user's gaze direction and intention.
[0057] The blinking detection and filtering module 32 is used to judge whether it is a blinking action by analyzing the frequency and amplitude characteristics of eye movement and perform filtering during the recognition process.
[0058] In this technical solution, the eye movement tracking sensor in the MR head-mounted display is responsible for collecting the original data of the user's eyes. These data include information such as the position of the eyes, the change in the size of the pupils, and the movement trajectory of the eyeballs, and are collected at a certain frequency, for example: 120 - 240 frames per second. Blinking is usually accompanied by the closing and opening of the eyes, and specific changes in parameters such as eye movement position and pupil size will occur. For example, when blinking, the pupil will be temporarily blocked by the eyelid, and the eye movement tracking sensor will detect a sharp decrease or even disappearance of the pupil area, and the eye movement position will also be relatively fixed at the position where the eyelid is closed. At the same time, the frequency and amplitude of blinking also have certain characteristics. The normal blinking frequency is generally about 10 - 20 times per minute, and the blinking amplitude is relatively stable. In some special cases, such as when the eyes are fatigued or stimulated, the blinking frequency may increase and the amplitude may become larger.
[0059] Specifically, it includes the following steps:
[0060] Previously, the differential algorithm is used to calculate the first-order differential velocity and the second-order differential acceleration of the eye movement position data to obtain the velocity and acceleration of the eye movement.
[0061] In this technical solution, when a blink occurs, the eye movement velocity will first drop rapidly (eyes closed) and then rise rapidly (eyes open), and the acceleration will also have a corresponding peak. These characteristics can be captured by calculating the eye movement velocity and acceleration.
[0062] Use the Canny edge detection image recognition algorithm to determine the pupil boundary, calculate the pupil area, and obtain the pupil area change;
[0063] This technical solution, through the recognition of pupil boundaries and area calculation, indicates a blinking action when the pupil area drops below a certain threshold in a short period of time and then recovers.
[0064] Eye movement velocity, acceleration, and pupil area change are used as feature inputs, and a support vector machine is used to determine whether it is a blink action.
[0065] Calibration is used in eye tracking to select preset scenes. If it is judged as a blink action, the blink action is filtered and feedback is prompted;
[0066] Specifically, in the scenario where eye tracking is used to select a virtual object, if a blink action is determined to be an eye movement, the system will not treat it as a selection operation for the object, thereby avoiding false triggering.
[0067] At the same time, when the system filters out a blink action, a small icon is displayed at the edge of the headset's field of view, indicating that the blink action just now was recognized and filtered, so that the user can understand the working status of the system and is used to indicate prompt feedback.
[0068] With the help of the above solution, through effective blink detection and filtering mechanism, unconscious eye movements such as blinking can be prevented from being misidentified as operation instructions, and the impact of environmental factors on eye tracking can be reduced.
[0069] The collaborative linkage module 4 is used for the gesture interaction control module 1, the voice interaction control module 2 and the eye movement interaction control module 3 to work together, and when the gesture interaction control module 1 or the voice interaction control module 2 or the eye movement interaction control module 3 performs priority control, the gesture interaction control module 1 or the voice interaction control module 2 or the eye movement interaction control module 3 is linked for collaborative control.
[0070] Specifically, in this technical solution, when a user touches a virtual object, the system detects information such as the position and force of the touch action through sensors in the MR headset, such as hand position sensors, force sensors, etc. After gesture recognition by the gesture interaction control module 1, it performs coordinated control with the voice interaction control module 2 and the eye movement interaction control module 3. Among them, for the voice interaction control module 2, it generates corresponding sounds according to the touch action and the characteristics of the object, and performs command control according to the user's voice signal. At the same time, for the eye movement interaction control module 3, it is used to track the user's eye movement according to the touch action and the characteristics of the object for coordinated control.
[0071] In summary, by means of the above technical solution of the present invention, the following effects can be achieved:
[0072] The present invention realizes the recognition of accurate complex gestures and the interactive recognition of command control through the gesture interaction control module 1, enabling users to interact with the virtual and real integrated environment more naturally and smoothly. For example, in an MR game scenario, if the gesture recognition is accurate, players can easily pick up virtual weapons, open virtual treasure chests, etc. just like in the real world. This seamless interaction experience will make users feel that they are truly part of the virtual world, greatly enhancing the sense of immersion. At the same time, when the voice interaction control module 2 and the eye movement interaction control module 3 achieve accurate interactive recognition, the user's attention can be more focused on the experience content itself, rather than being disturbed by frequent recognition errors. For example, users can select objects by natural eye fixation or quickly obtain information with voice commands, as if the virtual environment can understand their minds. The precise interactive recognition of the present invention reduces the fatigue of users caused by repeated operations or correcting incorrect operations. For the recognition of complex gestures and command control, users do not need to repeat the same action multiple times to achieve the goal, thus reducing the burden on the arm and hand muscles. Similarly, in improving the recognition of voice control commands in a noisy environment, accurate recognition means that users do not need to speak loudly or slowly repeatedly for the system to understand their intentions, thereby reducing the fatigue caused by voice interaction. For eye movement interaction control, filtering and command control are performed during the recognition of blink actions. Accurate recognition can prevent users from constantly adjusting their line of sight due to frequent false triggers, reducing eye strain and fatigue. For example, when using the headset for a long time to read virtual documents or observe virtual scenes, precise eye movement tracking can make users browse the content more comfortably.
[0073] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. After considering the disclosure in the specification and the embodiments, those skilled in the art will easily think of other implementation schemes of the present disclosure. This application aims to cover any variations, uses or adaptations of the present disclosure, which follow the general principles of the present disclosure and include the common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and the embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.
[0074] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. An MR headset intelligent control system, characterized in that include: A gesture interaction control module (1), a voice interaction control module (2) and an eye movement interaction control module (3), wherein; The gesture interaction control module (1) is used to configure a high-definition camera (11) and a depth sensor (12) to automatically adjust parameters in a complex environment and provide real-time gesture prompts and feedback, so as to perform recognition and command control of complex gestures; wherein the high-definition camera (11) is used to collect hand movements; and the depth sensor (12) is used to embed a gesture recognition algorithm to adapt to different hand movements, speeds and angles, so as to meet the requirements for recognition of complex gestures; The voice interaction control module (2) is used to filter and suppress noise in the environment, highlight the user's voice signal, and automatically adjust and optimize according to the user's voice characteristics through an embedded voice recognition engine to improve the recognition of voice control instructions in a noisy environment. It includes: a noise elimination module (21) and a voice enhancement optimization module (22), wherein the noise elimination module (21) is used to filter and suppress noise in the environment by using a microphone array and digital signal processing to highlight the user's voice signal; the voice enhancement optimization module (22) is used to embed a voice recognition engine and can automatically adjust and optimize according to the user's voice characteristics; The eye movement interaction control module (3) is used to track the eye movement of the user and to filter and control commands during the recognition process of blinking actions, and comprises: a calibration tracking module (31) and a blink detection filtering module (32), wherein the calibration tracking module (31) is used to perform eye movement calibration according to an eye movement tracking algorithm before using the head display; the blink detection filtering module (32) is used to determine whether it is a blinking action by analyzing the frequency and amplitude characteristics of the eye movement, and to filter it during the recognition process; The system further comprises: a collaborative linkage module (4) for enabling the gesture interaction control module (1), the voice interaction control module (2) and the eye movement interaction control module (3) to work in collaboration, and when the gesture interaction control module (1) or the voice interaction control module (2) or the eye movement interaction control module (3) performs priority control, the gesture interaction control module (1) or the voice interaction control module (2) or the eye movement interaction control module (3) is linked to perform collaborative control; the gesture interaction control module (1) further comprises: an environment calibration template (13) and a gesture guidance module (14), wherein; The environmental calibration template (13) is used to enable the head display to automatically adjust parameters in different usage environments to adapt to environmental changes, which at least includes: in an environment with too strong or too dark light, automatically adjusting the exposure parameters of the high-definition camera to ensure that hand movements can be clearly captured; The gesture guidance module (14) is used to provide real-time gesture prompts and feedback in the head display interface to help the user perform gesture operations correctly; The voice interaction control module (2) further comprises: a voice setting module (23) and a voice feedback module (24), wherein; The voice setting module (23) is used to provide voice templates and shortcut command settings, allowing users to customize voice commands; The voice feedback module (24) is configured to provide voice feedback after recognizing a user's voice command.
2. The MR headset intelligent control system according to claim 1, wherein The blink detection and filtering module (32) includes the following steps: Previously, a differential algorithm is used to calculate the first-order differential velocity and the second-order differential acceleration of the eye movement position data to obtain the velocity and acceleration of the eye movement. The Canny edge detection image recognition algorithm is used to determine the boundary of the pupil, calculate the pupil area, and obtain the change in the pupil area. Using the eye movement velocity, acceleration, and change in pupil area as feature inputs, a support vector machine is used to determine whether it is a blink action. Calibrated in eye movement tracking for selecting a preset scenario, if it is determined to be a blink action, the blink action is filtered and feedback is prompted.
Citation Information
Patent Citations
Virtual reality interaction system based on gesture, voice and sight line tracking recognition
CN109739353A