Multi-mode man-machine interaction control method and display and control integrated head-mounted device
Through the multimodal human-computer interaction control method, information from multiple sensory channels is collected and fused, human-computer interaction instructions are generated and actions are executed, and the problem that multimodal human-computer interaction in the prior art is not able to receive and fuse information from multiple sensory channels is solved, achieving a more efficient and consistent user experience.
Patent Information
- Application Number
- CN202311566567.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-22
- Publication Date
- 2025-05-23
AI Technical Summary
The existing multimodal human-computer interaction technology cannot effectively receive and fuse information between multiple sensory channels, resulting in a reduced consistency in user experience.
A multimodal human-computer interaction control method is provided. By responding to the environment monitoring instructions triggered by the user, the environment image collected by the unmanned device is obtained and displayed on the user's display page, multiple feedback actions of the user (such as gestures, electromyography and eye movements) are collected, and the information is fused using a preset multimodal human-computer interaction fusion algorithm, and human-computer interaction instructions are generated, and they are sent to the unmanned device to perform the actions.
It realizes the effective integration of multimodal information, improves the consistency of user perception, reduces manual intervention and operational needs, and improves operational efficiency and accuracy.
Smart Images

Figure CN120029439A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent wearable devices, and in particular to a multi-modal human-computer interaction control method and a display-control integrated head-mounted device. Background Art
[0002] Human-computer interaction technology builds a bridge for information exchange between humans and machines, enabling machines to better understand human intentions and help humans complete specific tasks faster and more efficiently. It is of great significance to improve production efficiency and economic benefits, enhance human experience, etc. Multimodal human-computer interaction technology can convert human's multiple senses into interactive instructions between machines, making the interaction between humans and machines more natural. Therefore, multimodal human-computer interaction technology is an important development direction and trend in human-computer interaction technology.
[0003] The currently developed multimodal human-computer interaction technology usually involves multiple interaction methods such as EEG, EMG, and gestures. However, there is currently a lack of methods to reasonably receive and fuse information between multiple sensory channels. The information obtained between multiple sensory channels is inconsistent, and it is impossible to integrate the information of all sensory channels to generate unified control instructions for operation, which in turn leads to reduced consistency of user experience. Summary of the invention
[0004] In view of the shortcomings of the above-mentioned prior art, the present invention provides a multimodal human-computer interaction control method and an integrated display and control head-mounted device, which solves the technical problem in the prior art that multimodal human-computer interaction cannot receive and integrate information between multiple sensory channels, thereby reducing the consistency of user experience.
[0005] On one hand, the present invention provides a multi-modal human-computer interaction control method, comprising:
[0006] In response to an environmental monitoring instruction triggered by a user, an environmental image collected by an unmanned device is acquired and displayed on a user's display page;
[0007] Collecting multiple feedback actions generated by the user based on the environment image, identifying and analyzing the multiple feedback actions, and generating multimodal action information;
[0008] Using a preset multi-modal human-computer interaction fusion algorithm, the multi-modal action information is fused and processed to generate a human-computer interaction instruction;
[0009] The human-computer interaction instruction is sent to the unmanned device, so that the unmanned device performs an action based on the human-computer interaction instruction.
[0010] Optionally, the multiple feedback actions include gesture actions, myoelectric actions and eye movements; the collecting the multiple feedback actions generated by the user based on the environment image, identifying and analyzing the multiple feedback actions, and generating multimodal action information includes:
[0011] Collecting the user's gestures, using the depth image to recognize the hand position information and the hand posture changes in the gestures, analyzing the hand position information and the hand posture changes, and obtaining gesture information;
[0012] Monitor the user's muscle activity, obtain the potential change of the muscle based on the muscle activity, collect the user's electromyographic action according to the potential change of the muscle and obtain the electromyographic action information;
[0013] Monitoring the user's eye movement state, collecting the user's eye movement based on the eye movement state, identifying the eye movement to obtain eye movement parameters, and generating eye movement information according to the eye movement parameters;
[0014] The gesture action information, the electromyographic action information and the eyeball action information are integrated to generate multimodal action information, wherein the multimodal action information is used to indicate the user's movement intention.
[0015] Optionally, the using a preset multimodal human-computer interaction fusion algorithm to perform fusion processing on the multimodal action information to generate a human-computer interaction instruction includes:
[0016] Preprocessing the multimodal action information;
[0017] Performing feature extraction on the preprocessed multimodal action information to obtain an action information feature vector, and performing representation learning on the action information feature vector to obtain an enhanced representation of the action information;
[0018] The enhanced representation of the action information is input into a preset multimodal human-computer interaction fusion algorithm for fusion processing to generate a human-computer interaction instruction.
[0019] Optionally, the training process of the multimodal human-computer interaction fusion algorithm includes:
[0020] Collecting multimodal sample data and user intention sample data, and preprocessing the multimodal sample data and the user intention sample data;
[0021] Performing feature extraction on the preprocessed multimodal sample data to obtain a sample feature vector, and performing representation learning on the sample feature vector to obtain an enhanced representation of the sample data;
[0022] Using a machine learning algorithm to model and classify the preprocessed user intent sample data to obtain a user intent modeling result;
[0023] The sample data enhanced representation, the user intention modeling results and the preset human-computer interaction effectiveness evaluation system are taken as input items and input into a preset machine learning model for training to generate a multimodal human-computer interaction fusion algorithm.
[0024] Optionally, sending the human-computer interaction instruction to the unmanned device so that the unmanned device performs an action based on the human-computer interaction instruction includes:
[0025] Sending the human-computer interaction instruction to the unmanned device in a preset data format, so that the unmanned device decodes the human-computer interaction instruction and generates a control signal according to the decoded human-computer interaction instruction;
[0026] Based on the control signal, the unmanned device performs an action and generates feedback information, and sends the feedback information to the display page of the user.
[0027] The multimodal human-computer interaction control method provided by the present invention first responds to the environmental monitoring instruction triggered by the user, obtains the environmental image collected by the unmanned equipment and displays it on the user's display page, collects multiple feedback actions generated by the user based on the environmental image, identifies and analyzes the multiple feedback actions, and generates multimodal action information. The present application collects the environmental image and displays it to the user, so that the user can understand and monitor the situation of the environment in real time, further identifies and analyzes the multiple feedback actions generated by the user based on the environmental image, and generates multimodal action information. The multimodal interaction allows the user to interact in different ways, with a more natural and flexible user experience; the multimodal action information is fused and processed by using a preset multimodal human-computer interaction fusion algorithm to generate a human-computer interaction instruction, and then autonomous decision-making and action control are performed according to the user's feedback and instructions. The fusion of the multimodal action information can effectively integrate the information obtained by the user from multiple sensory channels, thereby improving the consistency of user perception; the human-computer interaction instruction is sent to the unmanned equipment so that the unmanned equipment performs the action based on the human-computer interaction instruction, which can reduce the need for manual intervention and operation and improve the efficiency and accuracy of the operation.
[0028] Another aspect of the present invention provides a head-mounted device with integrated display and control, which is used to implement the multi-modal human-computer interaction control method as described above, and the device includes: a helmet body, a central processing unit, a wireless transmitter and receiver, a body-area Bluetooth receiver, and an eye tracking and display device;
[0029] The wireless transmitter and receiver is used to wirelessly communicate with the unmanned equipment, receive environmental images and send human-computer interaction instructions to achieve information interaction;
[0030] The body area Bluetooth receiver is used to receive multimodal motion information;
[0031] The eye tracking and display device is arranged on the front side of the helmet body and is arranged opposite to the eye position of the user. The eye tracking and display device is used to display the environment image and collect the eye movement of the user;
[0032] The central processor and the wireless transmitter and receiver are integrated into an integrated setting. The central processor is respectively connected to the body area Bluetooth receiver and the eye tracking and display device, and is used to receive the multimodal action information to generate the human-computer interaction instructions, and send them to the unmanned device through the wireless transmitter and receiver.
[0033] Optionally, a plurality of physical interfaces are provided on the helmet body;
[0034] The helmet body is connected to the central processor, the wireless transmitter and receiver, the body area Bluetooth receiver and the eye tracking and display device respectively through the physical interface.
[0035] Optionally, the eye tracking and display device includes a low-light camera and AR display glasses, and the low-light camera is integrated on the AR display glasses;
[0036] The low-light camera is used to collect the user's eye movements, generate eye movement information, and send the eye movement information to the central processor;
[0037] The AR display glasses are used to receive environmental images and display the environmental images to the user.
[0038] Optionally, the device includes a windproof mask;
[0039] The windproof mask is arranged on the front side of the helmet body and is arranged opposite to the user's face;
[0040] The windproof mask is rotatably connected to the inner cavity of the helmet body so that the windproof mask can be stored in the inner cavity of the helmet body.
[0041] Optionally, the device further comprises a data glove, on which a Bluetooth terminal is arranged, and the Bluetooth terminal is connected to the body area Bluetooth receiver via Bluetooth communication;
[0042] The data glove is used to collect the user's gestures, generate gesture information, and send the gesture information to the central processor through the Bluetooth terminal.
[0043] The display-control integrated head-mounted device provided by the present invention is light and easy to carry, and can be worn directly on the head, so that the user can achieve more natural and convenient multi-modal human-computer interaction without handheld devices; environmental images are collected through eye tracking and display devices, and displayed within the user's field of vision, thereby achieving efficient environmental perception and monitoring capabilities, and controlling unmanned equipment to perform actions through interactive instructions based on the perception results; the central processing unit can receive information fed back from multiple modules, and then use a preset multi-modal human-computer interaction fusion algorithm to fuse the received information, generate human-computer interaction instructions, and achieve intelligent interaction; through the synergistic effect of the helmet body, the central processing unit, the wireless transmitter and receiver, the body-area Bluetooth receiver, and the eye tracking and display devices, multi-modal human-computer interaction and unmanned equipment control are achieved, which has the advantages of convenience, efficiency, and flexibility.
[0044] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description, claims, and drawings.
[0045] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0047] Figure 1 A schematic diagram of a flow chart of a multi-modal human-computer interaction control method in an embodiment provided in the present application;
[0048] Figure 2 A schematic diagram of the principle of the training process of the multimodal human-computer interaction fusion algorithm in the multimodal human-computer interaction control method in an embodiment provided in the present application;
[0049] Figure 3 A schematic diagram of an application flow of a multi-modal human-computer interaction control method in an embodiment provided in the present application;
[0050] Figure 4 A schematic diagram of the structure of a head-mounted device with integrated display and control in one embodiment provided in the present application;
[0051] Figure 5 This is a schematic structural diagram of a head-mounted device with integrated display and control in another embodiment provided in the present application.
[0052] In the figure:
[0053] 1. Helmet body; 2. Central processing unit; 3. Body-area Bluetooth receiver; 4. Eye tracking and display equipment; 5. Windproof mask; 6. Data gloves; 601. Bluetooth terminal. DETAILED DESCRIPTION
[0054] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise" and the like indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the referred device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as limiting the present invention.
[0055] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0056] In the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", "connected", "fixed" and the like should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be an indirect connection through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0057] On the one hand, the present invention provides a multi-modal human-computer interaction control method, such as Figure 1 As shown, it specifically includes first responding to the environmental monitoring command triggered by the user, obtaining the environmental image collected by the unmanned equipment and displaying it on the user's display page, then collecting multiple feedback actions generated by the user based on the environmental image, identifying and analyzing the multiple feedback actions, generating multimodal action information, and then using a preset multimodal human-computer interaction fusion algorithm to fuse the multimodal action information to generate a human-computer interaction command, and finally sending the human-computer interaction command to the unmanned equipment, so that the unmanned equipment performs an action based on the human-computer interaction command.
[0058] The multimodal human-computer interaction control method provided by the present invention first responds to the environmental monitoring instruction triggered by the user, obtains the environmental image collected by the unmanned equipment and displays it on the user's display page, collects multiple feedback actions generated by the user based on the environmental image, identifies and analyzes the multiple feedback actions, and generates multimodal action information. The present application collects the environmental image and displays it to the user, so that the user can understand and monitor the situation of the environment in real time, further identifies and analyzes the multiple feedback actions generated by the user based on the environmental image, and generates multimodal action information. The multimodal interaction allows the user to interact in different ways, and has a more natural and flexible user experience; the multimodal action information is fused and processed by using a preset multimodal human-computer interaction fusion algorithm to generate a human-computer interaction instruction, and then autonomous decision-making and action control are performed according to the user's feedback and instructions. The fusion of the multimodal action information can effectively integrate the information obtained by the user from multiple sensory channels, thereby improving the consistency of user perception; the human-computer interaction instruction is sent to the unmanned equipment so that the unmanned equipment performs the action based on the human-computer interaction instruction, which can reduce the need for manual intervention and operation and improve the efficiency and accuracy of the operation.
[0059] Furthermore, the multiple feedback actions include gesture actions, electromyographic actions and eye movements; wherein, multiple feedback actions generated by the user based on the environmental image are collected, the multiple feedback actions are identified and analyzed, and multimodal action information is generated, specifically including first collecting the user's gesture actions, using the depth image to identify the hand position information and hand posture changes in the gesture actions, analyzing the hand position information and hand posture changes to obtain gesture action information; monitoring the user's muscle activity, obtaining the muscle potential changes based on the muscle activity, collecting the user's electromyographic actions according to the muscle potential changes and obtaining the electromyographic action information; then monitoring the user's eye movement state, collecting the user's eye movement based on the eye movement state, identifying the eye movement to obtain eye movement parameters, generating eye movement information according to the eye movement parameters, and finally integrating the gesture action information, electromyographic action information and eye movement information to generate multimodal action information, wherein the multimodal action information is used to indicate the user's movement intention.
[0060] Specifically, the present application provides a method for collecting gestures, electromyographic movements and eye movements and generating multimodal motion information. Specifically, gestures can be obtained by using devices such as depth cameras, RGB cameras, cameras and gesture recognition sensors to obtain user gesture information, and recognize gestures by monitoring the user's gestures, thereby further identifying the user's intentions. For example, a depth camera can use a depth image to recognize the position and posture changes of the hand, and then recognize gestures, while a gesture recognition sensor specifically recognizes and tracks gestures; electromyographic motion information usually comes from signals collected by electrodes attached to human muscles, so it is necessary to use electromyographic sensors or bioelectric signal amplifiers and other devices to collect it. Electromyographic sensors are usually attached to the body to detect muscle activity, and detect the contraction and relaxation of muscles by measuring the potential changes of muscles, so they can reflect the user's movement intentions and states; eye movement information can be obtained using eye tracking devices, and the user's focus of attention, viewing direction and other information can be identified by monitoring the movement of the eyeballs. Typical eye movement parameters include gaze position, gaze duration, gaze frequency, etc.
[0061] In this embodiment, feedback information from multiple aspects such as user gestures, electromyographic signals, and eye movements is collected, and the user's behavior is monitored and analyzed from multiple sensory channels to provide comprehensive feedback information, thereby helping the user to express his or her intentions more accurately, and avoiding the limitation of single feedback information, generating multimodal action information, and more comprehensively understanding the user's movement intentions, thereby achieving all-round human-computer interaction, allowing the user to operate more naturally, and improving the convenience and comfort of operation. Multimodal feedback technology can be applied to various human-computer interaction scenarios, such as virtual reality, smart homes, unmanned inspections, and other fields, thereby providing a richer and more intelligent application experience.
[0062] Furthermore, the multimodal action information is fused using a preset multimodal human-computer interaction fusion algorithm to generate human-computer interaction instructions, including first preprocessing the multimodal action information, then extracting features from the preprocessed multimodal action information to obtain a feature vector of the action information, and performing representation learning on the feature vector of the action information to obtain an enhanced representation of the action information, and finally inputting the enhanced representation of the action information into a preset multimodal human-computer interaction fusion algorithm for fusion processing to generate human-computer interaction instructions.
[0063] Specifically, the multimodal motion information is first preprocessed. The multimodal motion information specifically includes gesture motion information, electromyographic information, and eye motion information. The preprocessing process may include noise removal, data smoothing, coordinate system calibration, dimensionality reduction, and other operations to improve the accuracy and reliability of subsequent feature extraction and analysis; feature extraction is performed on the preprocessed multimodal motion information. The purpose of feature extraction is to convert the original data into a specific numerical representation to reflect the key features of the action. For gesture motion information, features such as hand position, hand posture changes, and gesture shape may be extracted; for electromyographic motion information, spectral features or time domain features of muscle potential changes may be extracted; for eye motion information, features such as eye movement trajectory and gaze point position may be extracted; representation learning is performed on the motion information feature vector obtained by feature extraction. Through representation learning, the potential structure and pattern in the data can be discovered, and a more compact and discriminative enhanced representation of motion information can be obtained; the multimodal human-computer interaction fusion algorithm can comprehensively utilize the information advantages of different modalities to improve the accuracy and reliability of interactive instructions.
[0064] In this embodiment, through the steps of preprocessing, feature extraction, representation learning and fusion processing, the accuracy and reliability of feedback information are effectively improved, the fusion of multimodal information is achieved, and the intelligence level of human-computer interaction is improved, thereby improving user experience and satisfaction.
[0065] Furthermore, the training process of the multimodal human-computer interaction fusion algorithm is as follows Figure 2 As shown, the method includes first collecting multimodal sample data and user intention sample data, preprocessing the multimodal sample data and user intention sample data, then performing feature extraction on the preprocessed multimodal sample data to obtain a sample feature vector, and performing representation learning on the sample feature vector to obtain an enhanced representation of the sample data, and then using a machine learning algorithm to model and classify the preprocessed user intention sample data to obtain a user intention modeling result, and finally using the sample data enhanced representation, the user intention modeling result and a preset human-computer interaction effectiveness evaluation system as input items, and inputting them into a preset machine learning model for training to generate a multimodal human-computer interaction fusion algorithm.
[0066] Specifically, the collected multimodal sample data and user intention sample data are first preprocessed, including data cleaning, noise filtering, normalization, data alignment and other operations. Then, useful features are extracted from the preprocessed multimodal data, and representation learning is performed. The feature extraction method can use principal component analysis, discrete wavelet transform, or use a deep learning model for end-to-end feature extraction and representation learning. Then, machine learning algorithms such as support vector machines, decision trees, and deep neural networks are used to model and classify user intentions. By analyzing and predicting user operation behaviors and interaction goals, the user's current intentions can be determined, which facilitates a better understanding of the user's behavior and intentions. And respond more accurately to user requirements; then enhance the representation of the samples by using unsupervised or supervised reinforcement learning algorithms, so as to obtain feature vectors that are more discriminative and distinguishable than the original data, improve the accuracy and robustness of the algorithm, and thus improve the performance of the algorithm in multimodal human-computer interaction scenarios; finally, the enhanced representation of sample data, user intention modeling results and the preset human-computer interaction effectiveness evaluation system are used as input items and input into the preset machine learning model for training to generate a multimodal human-computer interaction fusion algorithm. Specifically, classifiers, neural networks and other machine learning models can be used for training, and back-propagation algorithms and other technologies can be used for optimization and updating to improve the algorithm effect and performance.
[0067] In this embodiment, by extracting multimodal features, modeling user intent, and utilizing machine learning to train the model, this training process can improve the performance and effectiveness of the multimodal human-computer interaction fusion algorithm, enabling the algorithm to better understand user intent, accurately respond to user needs, and improve user experience and interaction effects.
[0068] Furthermore, the human-computer interaction instruction is sent to the unmanned device so that the unmanned device performs an action based on the human-computer interaction instruction, including first sending the human-computer interaction instruction to the unmanned device in a preset data format so that the unmanned device decodes and processes the human-computer interaction instruction, and generates a control signal based on the decoded human-computer interaction instruction, and then based on the control signal, the unmanned device performs the action and generates feedback information, and sends the feedback information to the user's display page.
[0069] In this embodiment, the human-computer interaction instruction is sent to the unmanned device, and the unmanned device itself can decode the human-computer interaction instruction to generate a control signal, and then perform an action based on the control signal, so that the unmanned device can autonomously complete the task guided by the user, and can complete the task more accurately and efficiently, wherein the unmanned device specifically includes drones and unmanned vehicles. This can avoid the influence of human factors on the execution of tasks, and the feedback information of the unmanned device is sent to the user's display page, so that the user can understand the progress of the task execution, so as to better grasp the progress and results of the task; and by sending the human-computer interaction instruction to the unmanned device in a preset data format through a wireless communication network, the interaction delay and communication cost can be effectively reduced, and the interaction efficiency can be improved, so that the human-computer interaction is more efficient and convenient.
[0070] The specific process of the multi-modal human-computer interaction control method provided by the present invention is as follows: Figure 3 As shown, first wait for the system to be prepared. After the preparation is completed, the system starts to collect the user's movements through multiple sensory channels, including gestures, electromyography and eye movements, and then obtains gesture action information, electromyography action information and eye movement information. The obtained multimodal action information is fused through a multimodal human-computer interaction algorithm to generate human-computer interaction instructions, which are sent to the unmanned equipment to control the unmanned equipment to perform actions and complete the entire control process.
[0071] Another aspect of the present invention provides a display-control integrated head-mounted device for implementing the above-mentioned multi-modal human-computer interaction control method, such as Figure 4 As shown, the device includes: a helmet body 1, a central processing unit 2, a wireless transmitter and receiver, a body area Bluetooth receiver 3 and an eye tracking and display device 4, wherein the wireless transmitter and receiver is used to wirelessly communicate with the unmanned device, for receiving environmental images and sending human-computer interaction instructions to realize information interaction, the body area Bluetooth receiver 3 is used to receive multimodal motion information, the eye tracking and display device 4 is arranged on the front side of the helmet body 1, and is arranged relative to the user's eye position, the eye tracking and display device 4 is used to display the environmental image and collect the user's eye movement, the central processing unit 2 and the wireless transmitter and receiver are integrated, the central processing unit 2 is respectively connected to the body area Bluetooth receiver 3 and the eye tracking and display device 4, for receiving multimodal motion information to generate human-computer interaction instructions, and sending them to the unmanned device through the wireless transmitter and receiver.
[0072] The display and control integrated head-mounted device provided by the present invention is light and easy to carry, and can be directly worn on the head. The user can achieve more natural and convenient multi-modal human-computer interaction without handheld devices; the collected environmental image is obtained through the eye tracking and display device 4, and displayed within the user's field of vision, to achieve efficient environmental perception and monitoring capabilities, and the unmanned equipment is controlled to perform actions through interactive instructions based on the perception results; the central processor 2 can receive information from multiple modules, and then use the preset multi-modal human-computer interaction fusion algorithm to fuse the received information, generate human-computer interaction instructions, and achieve intelligent interaction; through the synergy of the helmet body 1, the central processor 2, the wireless transmitter and receiver, the body domain Bluetooth receiver 3 and the eye tracking and display device 4, multi-modal human-computer interaction and unmanned equipment control are achieved, which has the advantages of convenience, efficiency, flexibility, etc. Among them, the wireless transmitter and receiver central processor 2 can be set on the top of the helmet body to obtain better communication effects, and can dissipate heat in time to improve the performance of the processor, and the body domain Bluetooth receiver 3 is usually set at the corresponding position of the ear of the helmet body 1 to balance the overall weight, and the specific position needs to be set according to the connected sensors and other equipment.
[0073] Furthermore, a plurality of physical interfaces are provided on the helmet body 1, wherein the helmet body 1 is connected to the central processing unit 2, the wireless transmitter and receiver, the body domain Bluetooth receiver 3 and the eye tracking and display device 4 respectively through the physical interfaces. In the present embodiment, the physical interface refers to the physical connection point for data transmission between a computer or other electronic device and an external device, and common physical interfaces on the helmet body 1 include a USB interface, an HDMI interface, an audio interface, an Ethernet interface, a Bluetooth interface and an infrared interface. The helmet body 1 provided in the present application is connected and transmits data with various devices through different types of physical interfaces, thereby realizing more functions and interaction modes, improving the application value and user experience of the helmet body 1, and with the help of a variety of physical interfaces, the central processing unit, the wireless transmitter and receiver, the body domain Bluetooth receiver 3 and the eye tracking and display device 4 can be quickly assembled on the helmet, and it is convenient to replace and repair, and has good integration and replaceability.
[0074] Furthermore, the eye tracking and display device 4 includes a low-light camera and AR display glasses. The low-light camera is integrated in the AR display glasses. The low-light camera is used to collect the user's eye movements, generate eye movement information, and send the eye movement information to the central processor 2. The AR display glasses are used to receive environmental images and display the environmental images to the user.
[0075] In this embodiment, the user's eye movement information is collected by a low-light camera, and the visual output can be presented at the position of the user's line of sight, which greatly improves the naturalness and operability of the user's interaction; the environmental image can be displayed to the user through AR display glasses, and the data can be displayed more intuitively by superimposing virtual information on the environmental image; the user's eye movement is tracked by the low-light camera, so the user's line of sight focus can be located more accurately, and the helmet body 1's perception of the user is improved; the low-light camera and AR display glasses are integrated into a design so that the structure occupies a smaller space and consumes lower power.
[0076] Further, such as Figure 5 As shown, the device includes a windproof mask 5, which is arranged on the front side of the helmet body 1 and opposite to the user's face. The windproof mask 5 is rotatably connected to the inner cavity of the helmet body 1 so that the windproof mask 5 can be stored in the internal cavity of the helmet body 1.
[0077] In this embodiment, the windproof mask 5 can block the damage to the eyes caused by external impurities such as wind, sand and rain, and protect the user's eyesight. The windproof mask 5 can be adapted to more weather and environmental conditions, such as rainy days and high wind speed conditions. The windproof mask 5 is arranged relative to the user's face, which can effectively reduce the interference of wind noise and other noises and improve the wearing comfort. Through the rotation connection between the windproof mask 5 and the inner cavity of the helmet body 1, the windproof mask 5 can be stored in the internal cavity of the helmet body 1, which is convenient for storage and carrying when not in use. In summary, the windproof mask 5 is rotatably connected to the inner cavity of the helmet body 1, which can bring better eye protection effect, enhance the applicability of the helmet body 1, improve wearing comfort and more convenient storage, thereby improving the convenience, safety and comfort of the use of the helmet, and providing users with a better use experience.
[0078] Further, such as Figure 5 As shown, the device also includes a data glove 6, on which a Bluetooth terminal 601 is provided. The Bluetooth terminal 601 is connected to the body area Bluetooth receiver 3 via Bluetooth communication. The data glove 6 is used to collect the user's gestures, generate gesture information, and send the gesture information to the central processor 2 via the Bluetooth terminal 601.
[0079] In this embodiment, the user's gestures are collected by the data glove 6, and the corresponding gesture information is generated. Through the communication between the Bluetooth terminal 601 and the body domain Bluetooth receiver 3, the real-time gestures can be captured and transmitted, accurately reflecting the changes in the user's gestures. The data glove 6 can recognize the user's gestures with high precision. Through the connection between the Bluetooth terminal 601 and the body domain Bluetooth receiver 3, the gesture information can be accurately transmitted to the central processor 2 for further analysis and processing; through the gesture capture and recognition of the data glove 6, the user can control the device or application through gestures to achieve a more natural and intuitive interaction method. For example, in a virtual reality environment, the selection, operation or navigation of objects can be completed through gestures; and because the data glove 6 uses Bluetooth wireless communication, it can achieve remote connection with the central processor 2, reducing the restraint of the device, and the user can move and operate more freely. In summary, the combination of the Bluetooth terminal 601 of the data glove 6 and the body domain Bluetooth receiver 3 can achieve the beneficial effects of real-time gesture capture, high-precision gesture recognition, enhanced interactivity, flexible portability and hand movement analysis, providing users with a richer and more natural interactive experience and data collection function.
[0080] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A multimodal human-computer interaction control method, It is characterized in that The method comprises: In response to an environmental monitoring instruction triggered by a user, an environmental image collected by an unmanned device is acquired and displayed on a user's display page; Collecting multiple feedback actions generated by the user based on the environment image, identifying and analyzing the multiple feedback actions, and generating multimodal action information; Using a preset multi-modal human-computer interaction fusion algorithm, the multi-modal action information is fused and processed to generate a human-computer interaction instruction; The human-computer interaction instruction is sent to the unmanned device, so that the unmanned device performs an action based on the human-computer interaction instruction.
2. The method according to claim 1, It is characterized in that The multiple feedback actions include gesture actions, electromyographic actions, and eyeball actions; the collecting multiple feedback actions generated by the user based on the environment image, identifying and analyzing the multiple feedback actions, and generating multimodal action information includes: Collecting the user's gestures, using the depth image to recognize the hand position information and the hand posture changes in the gestures, analyzing the hand position information and the hand posture changes, and obtaining gesture information; Monitor the user's muscle activity, obtain the potential change of the muscle based on the muscle activity, collect the user's electromyographic action according to the potential change of the muscle and obtain the electromyographic action information; Monitoring the user's eye movement state, collecting the user's eye movement based on the eye movement state, identifying the eye movement to obtain eye movement parameters, and generating eye movement information according to the eye movement parameters; The gesture action information, the electromyographic action information and the eyeball action information are integrated to generate multimodal action information, wherein the multimodal action information is used to indicate the user's movement intention.
3. The method according to claim 1, It is characterized in that The method of using a preset multi-modal human-computer interaction fusion algorithm to fuse the multi-modal action information and generate a human-computer interaction instruction includes: Preprocessing the multimodal action information; Performing feature extraction on the preprocessed multimodal action information to obtain an action information feature vector, and performing representation learning on the action information feature vector to obtain an enhanced representation of the action information; The enhanced representation of the action information is input into a preset multimodal human-computer interaction fusion algorithm for fusion processing to generate a human-computer interaction instruction.
4. The method according to claim 1, It is characterized in that The training process of the multimodal human-computer interaction fusion algorithm includes: Collecting multimodal sample data and user intention sample data, and preprocessing the multimodal sample data and the user intention sample data; Performing feature extraction on the preprocessed multimodal sample data to obtain a sample feature vector, and performing representation learning on the sample feature vector to obtain an enhanced representation of the sample data; Using a machine learning algorithm to model and classify the preprocessed user intent sample data to obtain a user intent modeling result; The sample data enhanced representation, the user intention modeling results and the preset human-computer interaction effectiveness evaluation system are taken as input items and input into a preset machine learning model for training to generate a multimodal human-computer interaction fusion algorithm.
5. The method according to claim 1, It is characterized in that The step of sending the human-computer interaction instruction to the unmanned device so that the unmanned device performs an action based on the human-computer interaction instruction includes: Sending the human-computer interaction instruction to the unmanned device in a preset data format, so that the unmanned device decodes the human-computer interaction instruction and generates a control signal according to the decoded human-computer interaction instruction; Based on the control signal, the unmanned device performs an action and generates feedback information, and sends the feedback information to the display page of the user.
6. A head-mounted device with integrated display and control, It is characterized in that Used to implement the multimodal human-computer interaction control method as claimed in any one of claims 1 to 5, the device comprises: a helmet body (1), a central processing unit (2), a wireless transmitter and receiver, a body area Bluetooth receiver (3) and an eye tracking and display device (4); The wireless transmitter and receiver is used to wirelessly communicate with the unmanned equipment, receive environmental images and send human-computer interaction instructions to achieve information interaction; The body area Bluetooth receiver (3) is used to receive multi-modal action information; The eye tracking and display device (4) is arranged on the front side of the helmet body (1) and is arranged relative to the position of the user's eyeballs. The eye tracking and display device (4) is used to display environmental images and collect the user's eyeball movements; The central processor (2) and the wireless transmitter and receiver are integrated into an integrated configuration. The central processor (2) is respectively connected to the body area Bluetooth receiver (3) and the eye tracking and display device (4) to receive the multimodal action information to generate the human-computer interaction instructions, and send them to the unmanned device via the wireless transmitter and receiver.
7. The display-control integrated head-mounted device according to claim 6, It is characterized in that The helmet body (1) is provided with a plurality of physical interfaces; The helmet body (1) is connected to the central processor (2), the wireless transmitter and receiver, the body area Bluetooth receiver (3) and the eye tracking and display device (4) respectively through the physical interface.
8. The display-control integrated head-mounted device according to claim 6, It is characterized in that The eye tracking and display device (4) comprises a low-light camera and AR display glasses, wherein the low-light camera is integrated on the AR display glasses; The low-light camera is used to collect the user's eye movements, generate eye movement information, and send the eye movement information to the central processor (2); The AR display glasses are used to receive environmental images and display the environmental images to the user.
9. The display-control integrated head-mounted device according to claim 6, It is characterized in that The device comprises a windproof mask (5); The windproof mask (5) is arranged on the front side of the helmet body (1) and is arranged opposite to the face of the user; The windproof mask (5) is rotatably connected to the inner cavity of the helmet body (1), so that the windproof mask (5) can be stored in the inner cavity of the helmet body (1).
10. The display-control integrated head-mounted device according to claim 6, It is characterized in that The device further comprises a data glove (6), on which a Bluetooth terminal (601) is arranged, and the Bluetooth terminal (601) is connected to the body area Bluetooth receiver (3) via Bluetooth communication; The data glove (6) is used to collect the user's gesture movements, generate gesture movement information, and send the gesture movement information to the central processor (2) via the Bluetooth terminal (601).