Multi-modal human-computer interface equipment, control method thereof and related equipment

By using multimodal interaction methods, information is acquired through cameras, fingerprints, and remote guidance modules to generate control commands, solving the problem of low operating efficiency of industrial HMI equipment in dusty or oily environments and achieving efficient and safe equipment operation.

CN120891920APending Publication Date: 2025-11-04XIAN HUICHUAN TECHNOLOGY R&D CENTER CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510996214.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

Existing industrial HMI equipment is prone to malfunction in dusty or oily environments, and the single interaction method leads to low operating efficiency. Furthermore, it is slow to respond or ineffective when operating with gloves, making it impossible to achieve rapid batch or complex command input.

Method used

It adopts a multimodal interaction method, including a camera module to acquire user image information, a fingerprint module to acquire fingerprint information, and a remote guidance module to acquire remote guidance information, and generates control commands through a processor to operate the device.

Benefits of technology

It improves equipment operating efficiency, supports multiple interaction methods, reduces on-site travel time, enhances safety and ease of operation, and avoids safety hazards introduced by external equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120891920A_ABST
    Figure CN120891920A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-mode man-machine interface device, a control method thereof and a related device, and relates to the technical field of industrial control, the multi-mode man-machine interface device comprises an information obtaining device and a processor device, the information obtaining device responds to interaction operation of a user, obtains interaction information according to the interaction operation, and sends the interaction information to the processor device; and the processor device generates a control instruction according to the interaction information, so that the human-computer interface equipment performs equipment operation according to the control instruction. A user can interact with the multi-modal human-computer interface device through different interactive operations, the information acquisition device acquires different interactive information according to the different interactive operations, the processor device generates corresponding control instructions according to the different interactive information, and the human-computer interface device can perform device operation according to the control instructions. Therefore, equipment operation is carried out in a non-single interaction mode; the technical problem of low equipment operation efficiency can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of industrial control, and in particular to a multi-modal human-machine interface device, a control method thereof, and related devices. BACKGROUND

[0002] At present, the operation of an industrial HMI (Human Machine Interface) product is usually realized through a single interaction mode of touching a touch screen or a mechanical button.

[0003] However, the touch screen is prone to malfunction in an industrial site environment full of dust or oil stains, resulting in failure to work normally or misoperation, and the touch screen also has the problem of slow or even ineffective response when a user operates the touch screen while wearing gloves; and the manual pressing of mechanical buttons one by one cannot realize the input of fast batch or complex instructions; that is, the single interaction mode reduces the operation efficiency of the multi-modal human-machine interface device. SUMMARY

[0004] The main purpose of the present application is to provide a multi-modal human-machine interface device, a control method thereof, and related devices, aiming to solve the technical problem of low device operation efficiency.

[0005] To achieve the above purpose, the present application provides a multi-modal human-machine interface device, which comprises:

[0006] An information acquisition device is configured to acquire interaction information according to an interactive operation of a user in response to the interactive operation; the information acquisition device comprises a camera module, a fingerprint module, and a remote guidance module; the camera module is configured to acquire user image information, the fingerprint module is configured to acquire user fingerprint information, and the remote guidance module is configured to acquire remote guidance information; the interaction information comprises at least one of the user image information, the fingerprint information, and the remote guidance information;

[0007] A processor device is configured to generate a control instruction according to the interaction information, so that the human-machine interface device performs device operation according to the control instruction.

[0008] In an embodiment, the user image information comprises user gesture information, and the control instruction comprises a system operation instruction; the camera module is configured to acquire the user gesture information; and the processor device is configured to generate a system operation instruction according to the user gesture information, so that the human-machine interface device performs device system operation according to the system operation instruction.

[0009] In an embodiment, the user image information further comprises user face information; and the processor device is configured to generate a user permission verification instruction according to the fingerprint information and the face information, so that the human-computer interface device verifies the operation permission of the user according to the user permission verification instruction.

[0010] In an embodiment, the remote guidance module comprises:

[0011] a wireless network module configured to establish a network connection with a remote device;

[0012] an audio / video module configured to acquire an audio / video signal;

[0013] The processor device is configured to generate a remote conference operation instruction according to the audio / video signal in response to a conference request initiated by the remote device, so that the human-computer interface device establishes a remote conference mode with the remote device.

[0014] To achieve the above object, the present application further provides a control method of a multi-modal human-computer interface device, which is applied to the multi-modal human-computer interface device as described in any one of the above embodiments, and comprises:

[0015] In response to an interactive operation of a user, interactive information is acquired according to the interactive operation, wherein the interactive information comprises at least one of user image information, fingerprint information and remote guidance information;

[0016] A control instruction is generated according to the interactive information, so that the human-computer interface device performs device operation according to the control instruction.

[0017] In an embodiment, the user image information comprises user gesture information, the control instruction comprises a system operation instruction, and the step of generating a control instruction according to the interactive information, so that the human-computer interface device performs device operation according to the control instruction comprises:

[0018] determining a system operation instruction corresponding to the user gesture information;

[0019] performing system operation on the human-computer interface device according to the system operation instruction.

[0020] In an embodiment, the user image information further comprises user face information, and the step of generating a control instruction according to the interactive information, so that the human-computer interface device performs device operation according to the control instruction comprises:

[0021] comparing the user image information with preset target image information in terms of image similarity;

[0022] When the image similarity is greater than or equal to a target image similarity, comparing the user fingerprint information with preset target fingerprint information in terms of a fingerprint similarity;

[0023] When the fingerprint similarity is greater than or equal to a target fingerprint similarity, generating the user permission verification instruction;

[0024] Verifying the operation permission of the user according to the user permission verification instruction.

[0025] In an embodiment, the step of acquiring interaction information according to the interactive operation of the user in response to the interactive operation of the user comprises:

[0026] In response to a conference request initiated by a remote device, establishing a network connection with the remote device;

[0027] Acquiring remote guidance information according to the conference request, wherein the remote guidance information comprises audio and video information;

[0028] Generating a remote conference operation instruction according to the audio and video information;

[0029] Establishing a remote conference mode with the remote device according to the remote conference operation instruction.

[0030] In addition, to achieve the above object, the present application also provides an electronic device, which comprises a memory, a processor and a computer program stored in the memory and executable on the processor, and the computer program is configured to implement the steps of the control method of the multi-modal human-computer interface device as described above.

[0031] In addition, to achieve the above object, the present application also provides a storage medium, which is a computer readable storage medium, and the storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of the control method of the multi-modal human-computer interface device as described above.

[0032] In addition, to achieve the above object, the present application also provides a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the steps of the control method of the multi-modal human-computer interface device as described above.

[0033] The one or more technical solutions provided by the present application have at least the following technical effects:

[0034] The multi-modal human-computer interface device of the present application comprises an information acquisition device and a processor device, the information acquisition device acquires interaction information according to an interactive operation of a user in response to the interactive operation of the user, and the processor device generates a control instruction according to the interaction information, so that the human-computer interface device performs device operation according to the control instruction.

[0035] As the information acquisition device includes a camera module, a fingerprint module and a remote guidance module, the camera module is configured to acquire user image information, the fingerprint module is configured to acquire user fingerprint information, and the remote guidance module is configured to acquire remote guidance information, and the interactive information includes at least one of the user image information, the fingerprint information and the remote guidance information; that is, the user can realize interaction with the multi-modal human-computer interface device through different interactive operations, the information acquisition device acquires different interactive information according to different interactive operations, the processor device generates corresponding control instructions according to different interactive information, and the human-computer interface device can perform device operation according to the control instructions, so as to realize device operation through non-single interactive mode. Compared with the related art, only single interactive mode of touch control screen or mechanical button can be used to realize device operation, and the device operation efficiency can be improved. BRIEF DESCRIPTION OF DRAWINGS

[0036] The accompanying drawings, which are incorporated herein and constitute part of the specification, illustrate embodiments consistent with the application and, together with the description, serve to explain the principles of the application.

[0037] In order to more clearly explain the technical solutions of the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced as follows. Obviously, those skilled in the art can obtain other drawings according to these drawings without any creative effort.

[0038] Figure 1 A functional block diagram of the multi-modal human-computer interface device of the present application;

[0039] Figure 2 A flowchart of the control method of the multi-modal human-computer interface device of the present application;

[0040] Figure 3 A device structure diagram of the hardware running environment involved in the control method of the multi-modal human-computer interface device in the embodiments of the present application.

[0041] The object implementation, functional characteristics and advantages of the present application will be further explained with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION

[0042] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application, and are not used to limit the present application.

[0043] In order to better understand the technical solutions of the present application, the following will be described in detail in combination with the drawings and specific embodiments of the present application.

[0044] Embodiment one

[0045] The embodiment provides a multi-modal human-computer interface device, which comprises an information acquisition device and a processor device.

[0046] Since a single interaction mode of touching a touch screen or a mechanical button is used in the prior art, the operation of an industrial HMI (Human Machine Interface) product is realized. However, the touch screen is prone to failure in a dusty or oily industrial site environment, thereby causing normal operation to be impossible or misoperation to occur, and when a user wears gloves to operate the touch screen, the touch screen is also prone to slow response or even invalidity. In addition, the manual pressing mode of mechanical buttons cannot realize the input of fast batch or complex instructions, thereby reducing the operation efficiency of the multi-modal human-computer interface device.

[0047] The information acquisition device comprises a camera module, a fingerprint module and a remote guidance module.

[0048] Specifically, the camera module can comprise a camera, an image sensor, a control circuit and the like, and is used to acquire user image information, which can be body feature information or action feature information of the user or the like. The user image information acquired by the camera module can be a user image feature vector extracted by an image processing algorithm.

[0049] The fingerprint module device can comprise a capacitive fingerprint sensor, and is used to collect fingerprint information. The user fingerprint information collected by the capacitive fingerprint sensor can be that the user presses a finger on the surface of the capacitive fingerprint sensor, the capacitive unit array is scanned row by row, the capacitance value of each pixel point is measured, and the capacitance value is mapped into a gray-scale image (ridge is highlighted and valley is dark area). An algorithm is used to locate the detail feature points (such as bifurcation points, end points and isolated points) of the fingerprint, and usually 50-100 detail feature points are extracted, which contain position, direction, type and the like.

[0050] Since the multi-modal human-computer interface device is usually arranged in a field industrial scene, and special field industrial scenes (for example, a semiconductor production scene or a scene requiring secrecy) do not allow a user to bring a communication device (for example, a mobile phone, a computer or the like capable of communication), if the user needs remote guidance, the user can only leave the field industrial scene to perform the guidance and then return to the field industrial scene to operate the device, which greatly reduces the work efficiency.

[0051] To solve the problem, the remote guidance module is used to acquire remote guidance information, which can be sent to the local by a remote device. The multi-modal human-computer interface device can further comprise a microphone and a touch screen, and the remote guidance information can be displayed through the touch screen or played through the microphone.

[0052] The user can select to trigger the interactive operation of the human-computer interface device through a non-single mode such as a camera module, a fingerprint module, and a remote guidance module, for example, triggering an image acquisition operation through the camera module, triggering a fingerprint acquisition operation through the fingerprint module, triggering a remote guidance information acquisition operation through the remote guidance module, and the like. The information acquisition device acquires interactive information according to the interactive operation in response to the interactive operation of the user, at least one of user image information, fingerprint information, and remote guidance information.

[0053] The processor device is configured to generate a control instruction according to the interactive information, and the human-computer interface device can perform device operation according to the control instruction. The control instruction includes a device operation instruction, and the device operation instruction includes a device wake-up instruction, an authority verification instruction, a display instruction, or a playing instruction, and the like.

[0054] For example, the processor device generates a device wake-up instruction according to the user image information to wake up the human-computer interface device, the processor device generates an authority verification instruction according to the fingerprint information to start the authority verification function of the human-computer interface device, and the processor device generates a display instruction or a playing instruction according to the remote guidance information to display or play the remote guidance information through the touch screen or the microphone for the user to view.

[0055] Specifically, the remote guidance module includes a wireless network module, the wireless network module is configured to establish a network connection with a remote device, and an audio and video module, the audio and video module is configured to acquire an audio and video signal. The processor device is configured to generate a remote conference operation instruction according to the audio and video signal to make the human-computer interface device establish a remote conference mode with the remote device in response to a conference request initiated by the remote device.

[0056] Specifically, the wireless network module can be a WIFI module, a 4G module, a 5G module, and the like, allowing the device to access an existing wireless network environment or connect to the Internet through a mobile hotspot or the like. The audio and video module includes a microphone and a loudspeaker, and the audio input can be realized through the loudspeaker. The processor device is further configured to generate a remote conference operation instruction according to the audio and video signal to establish a remote conference mode with the remote device in response to a conference request initiated by the remote device. The touch screen is further configured to display conference content.

[0057] The remote device can be a remote mobile phone, a computer, a video conference terminal, and the like. In the above manner, the user can control the human-computer interface device and the remote device to establish a remote conference mode through voice interaction. The on-site personnel may be operating the device with both hands and cannot operate the touch screen or the keys. The human-computer interface device of the embodiment supports voice wake-up and voice control, for example, “start the conference”, “call Zhang Gong”, and the like, to realize hands-free operation and improve the operation convenience and safety.

[0058] In addition, remote video communication and remote control can be realized through the remote conference mode.

[0059] It should be noted that the information acquisition device can access the mainboard through an interface, and the processor device is also integrated on the mainboard. Therefore, the information acquisition device can send the collected interaction information to the processor device through the interface. Referring to Figure 1 , the camera, microphone, mechanical button and touch screen can access the mainboard through an interface, and different cameras, microphones, mechanical buttons and touch screens support different interfaces. The functional blocks can be introduced separately, and the functional blocks include a multi-media interface (Multi-Media Interface) block, a communication protocol block, and an audio / video block.

[0060] For example, the multi-media interface block (Multi-Media Interface) includes an LVDS (Low-Voltage Differential Signaling) interface or an MIPI-CSI_RX0 4Lane (MIPIDSI (Mobile Industry Processor Interface Display Serial Interface) is a mobile industry processor interface display serial interface, TX represents a sending end, and 4Lane represents four data channels for parallel data transmission) interface, an MIPI-CSI_RX1 4Lane (reserved) interface, which is another MIPI CSI receiving end port, and is used to expand the camera function; the multi-media interface block accesses the camera (camera module) and the touch screen (display / touch block) through different interfaces.

[0061] The communication protocol block (Connectivity) includes I2C, SPI, UART and various protocols. Among them, I2C includes I2C1_M0, I2C4_M0, I2C2 / 3 / 5 / 6M0; I2C1_M0 is used to connect Camera 1, EXT IO, LVDS_CTP, RTC, etc., I2C4_M0 is used to reserve connection of Camera 2, supports PCIe2.0, and I2C2 / 3 / 5 / 6M0 are all in the occupied state; SPI includes SPI0_M0 / SPI1_M1, SPI2_M0, SPI0_M0 / SPI1_M1 is in the occupied state, and is used for TF burning, SPI2_M0 is used to realize RFID related functions; UART includes UART8.

[0062] The wireless network module (such as a WIFI module) can establish a connection through an SDIO protocol and access through the interface of the multi-media interface block.

[0063] An audio video (Audio Video) block includes a MIC (MIC1_INL / MIC2_INR (B25)) and a SPK (SPKOUT_P / SPKOUT_N), the MIC is used to access a microphone, and the SPK is used to access a speaker.

[0064] In the embodiment, the user can realize interaction with the multi-modal man-machine interface device through different interaction operations, the information acquisition device acquires different interaction information according to different interaction operations, the processor device generates corresponding control instructions according to different interaction information, and the man-machine interface device can perform device operation according to the control instructions, so that device operation is realized through a non-single interaction mode. Compared with the related art, device operation can only be realized through a single interaction mode of touching a touch screen or a mechanical button, so that the device operation efficiency is improved. Remote guidance can be obtained without leaving the scene, the round trip time is reduced, and the speed of troubleshooting and device debugging is improved. The safety control of the on-site environment is maintained, and the safety hidden danger that may be caused by the introduction of external equipment is avoided.

[0065] Embodiment two

[0066] In order to further improve the device operation efficiency, the user image information of the embodiment includes user gesture information, and the control instruction includes a system operation instruction. The camera module is used to acquire the user gesture information. The processor device is used to generate the system operation instruction according to the user gesture information, so that the man-machine interface device performs device system operation according to the system operation instruction. That is, the user can realize device system operation through gesture interaction.

[0067] The user gesture information includes hand image, continuous multiple frames of hand image or hand motion video, etc. The control intention of the user can be determined according to the user gesture information, so that the corresponding system operation instruction is generated according to the control intention. The man-machine interface device sends the system operation instruction to the industrial control device, and realizes real-time control of the industrial control device.

[0068] Specifically, the system operation instruction includes an opening instruction, a stop instruction and the like of the industrial control device. The user can realize convenient and efficient device system operation through different gesture operations.

[0069] In addition, the zoom instruction and the like for controlling the display content of the touch screen can also be generated according to the gesture information.

[0070] Due to the more actions of the operating personnel in the industrial environment, the gesture operation is more likely to cause mis-touch. The information acquisition device can respond to the interactive operation of the user, determine the current operation mode according to the interactive operation, and then acquire the interactive information according to the current operation mode. The way of determining the current operation mode according to the interactive operation can be that when the interactive operation is a selection operation of the operation mode, the target operation mode is determined, and the target operation mode is entered, that is, the current operation mode is determined.

[0071] The embodiment that the information acquisition device acquires the interactive information according to the current operation mode can be that only when the interactive information corresponding to the current operation mode is acquired, a control instruction is further generated, so as to avoid interference caused by other interactive information.

[0072] For example, if the current operation mode is the wake-up mode, the device can be woken up only when the voice instruction is acquired. If the current operation mode is the fingerprint verification mode, the permission verification can be performed only when the user fingerprint information is acquired. If the current operation mode is the fingerprint verification mode, the corresponding gesture control instruction can be generated only when the gesture information is acquired.

[0073] It can be understood that the multi-modal human-computer interface device of the embodiment supports multiple operation modes, and different device operations can be performed according to different interactive information acquired by the information acquisition device in different operation modes, so as to realize non-single mode interaction. Compared with the single interaction mode of touching the touch screen or the mechanical button in the related art, the operation of the multi-modal human-computer interface device is realized. The embodiment can improve the device operation efficiency and solve the technical problem of low device operation efficiency.

[0074] Since the operation permission verification of the multi-modal human-computer interface device currently usually depends on single password input, identity recognition is performed according to the password, and thus the permission verification is realized. This way is easy to be maliciously cracked, which has high risk and poor security.

[0075] In order to improve the security of the device operation, the user image information further includes user face information. The processor device is configured to generate a user permission verification instruction according to the fingerprint information and the face information, so that the human-computer interface device verifies the operation permission of the user according to the user permission verification instruction.

[0076] The embodiment verifies the permission according to the multi-modal information, which is more difficult to be forged than the single-modal information verification, because the fingerprint information and the face information need to be acquired at the same time to pass the verification, which improves the security. At the same time, even if an error occurs in one modal information (such as a blurred fingerprint), the other modal information (such as face information) can still assist in judgment, which improves the overall recognition accuracy.

[0077] Specifically, the face information of the user can be collected by the camera module, the fingerprint information of the user can be acquired by the capacitive fingerprint sensor, the user permission verification instruction can be generated according to the face information and the fingerprint information, and the human-computer interface device can verify the operation permission of the user according to the user permission verification instruction.

[0078] The embodiment adopts the dual identity verification mechanism of face recognition and fingerprint recognition, and significantly improves the security level of device authorization.

[0079] It should be noted that voice information can also be acquired by a microphone, and retina information can be acquired by a retina scanner. The permission verification can be performed according to the voice information or the retina information. The permission verification according to the voice information can be performed by analyzing the voice characteristics (such as tone, speed, frequency, formant, etc.) of the user to identify the identity of the user, which is suitable for scenarios where it is inconvenient to touch with gloves, or in scenarios where it is inconvenient to use a camera or a fingerprint module, or in an identity confirmation link in remote collaboration guidance. The identity recognition can be performed by scanning the retinal blood vessel distribution pattern in the back of the eye by near-infrared light. The permission verification according to the retina information can be performed by scanning the retinal blood vessel distribution pattern in the back of the eye by near-infrared light.

[0080] Different user image information can be selected according to actual environment or security requirements for combined permission verification.

[0081] In the embodiment, the multi-modal human-computer interface device has a permission verification function based on multi-modal information. Compared with the traditional operation permission verification of the multi-modal human-computer interface device, which usually relies on a single password input and performs identity recognition according to the password to achieve permission verification, the permission verification based on multi-modal information in the embodiment is not easy to be maliciously cracked, and the security is improved.

[0082] Embodiment three

[0083] The embodiment provides a control method of a multi-modal human-computer interface device, which is applied to the multi-modal human-computer interface device in Embodiments One and Two. The same or similar contents as Embodiments One and Two can be referred to the foregoing description, and will not be described hereinafter. For details, refer to Figure 2 , Figure 2 The figure is a flowchart of the control method of the multi-modal human-computer interface device. In the embodiment, the control method of the multi-modal human-computer interface device includes steps S10-S20.

[0084] Step S10, in response to the interactive operation of the user, interactive information is acquired according to the interactive operation, wherein the interactive information includes at least one of user image information, fingerprint information and remote guidance information.

[0085] It should be noted that in order to solve the problem that the touch screen is prone to malfunction in an industrial site environment full of dust or oil stains, resulting in failure to work normally or misoperation, and when a user wears gloves to operate the touch screen, the touch screen will also have the problem of slow reaction or even invalidity, the user can realize interaction with the human-computer interface device through image input, fingerprint input, remote control and other ways, realize multi-modal interaction mode, and the user can complete the operation without touching the device, avoid pollution or safety risk, and support interaction in special environments such as wearing gloves and wearing protective clothing.

[0086] In order to improve the operation flexibility, the multi-modal human-computer interface device can determine the current operation mode according to the user behavior or the permission state.

[0087] Specifically, the implementation of determining the current operation mode can be:

[0088] In response to the wake-up operation, the verification information is obtained; according to the verification information, the permission verification is performed; when the permission verification is passed, the device operation mode is entered, and the current operation mode is determined according to the interaction information obtained in the device operation mode.

[0089] It can be understood that when the user triggers the human-computer interface device to wake up through a mechanical button, a gesture action or a voice instruction, the human-computer interface device starts an identity verification process; wherein the wake-up operation can be pressing the power / confirmation button, gesture close sensor (such as radar detection) or speaking a preset wake-up word.

[0090] Further, after the permission verification is successful, the device operation mode is entered, that is, the human-computer interface device enters an operable state; the multi-modal human-computer interface device continues to collect other interaction information of the user; according to the interaction information obtained in the device operation mode, the current operation mode is determined, that is, the type of operation (such as debugging mode, parameter setting mode, emergency stop, etc.) that the user currently wants to perform is determined, so as to avoid misoperation due to different interaction information in different operation modes.

[0091] Step S20, generating a control instruction according to the interaction information, so that the human-computer interface device performs device operation according to the control instruction.

[0092] Further, the human-computer interface device can generate a control instruction according to the interaction information, so that the human-computer interface device performs device operation according to the control instruction. Specifically, the way of performing device operation according to the interaction information can be to convert the interaction information into a control instruction, according to the type of the control instruction, to determine to execute the control instruction or send the control instruction to the industrial device.

[0093] For example, when the interaction type is determined as gesture sliding according to the interaction information, the interface switching instruction can be executed according to the type of the control instruction and the current operation mode (such as the interface display mode). When the interaction type is determined as the voice command "start" according to the interaction information, the start instruction can be sent to the industrial equipment according to the type of the control instruction and the current operation mode (such as the industrial equipment control mode), so as to control the industrial equipment.

[0094] Specifically, the human-computer interface device can generate a control instruction according to the interaction information required in the current operation mode. The type of the interaction information required to be obtained is determined according to different operation modes.

[0095] For example, in the "voice control mode", the microphone is enabled, and the voice information can be obtained. In the remote collaboration mode, the microphone, the speaker, the camera and the like are enabled, the voice information, the output audio, the video pictures captured and the like can be obtained.

[0096] Further, different device operations can be performed according to the interaction information obtained in different operation modes, so as to improve the device operation efficiency and flexibility, and avoid misoperation.

[0097] In addition, in order to avoid interference caused by other interaction information after entering a certain operation control mode (for example, after entering the gesture operation mode, the voice information can be obtained, thereby causing interference), when there are at least two types of interaction information, the device operation mode according to the interaction information can be: determining the interaction information to be responded according to the response priority of the at least two types of interaction information; and generating a control instruction according to the interaction information to be responded.

[0098] For example, the response priorities of different interaction information can be sorted in advance, or the response priorities of different interaction information can be determined according to different operation modes; for example, in the gesture control mode, the response priority of the gesture information is higher than that of the voice information, in the voice control mode, the response priority of the voice information is higher than that of the gesture information, and the like.

[0099] In order to avoid interference or misoperation, the interaction information to be responded can be determined according to the response priority of the at least two types of interaction information, and the control instruction can be generated according to the interaction information to be responded, so as to realize fast response.

[0100] Embodiment Four

[0101] In this embodiment, the user image information includes user gesture information, the control instruction includes a system operation instruction, and the implementation manner in which the human-computer interface device generates a control instruction according to the interaction information, so that the human-computer interface device performs device operation according to the control instruction can be:

[0102] determine a system operation instruction corresponding to the user gesture information; and perform system operation on the human-machine interface device according to the system operation instruction.

[0103] Specifically, different user gesture information corresponds to different system operation instructions. The gesture feature corresponding to the user gesture information can be determined according to an image recognition algorithm, the operation intention corresponding to the gesture can be determined according to the gesture feature, the corresponding system operation instruction can be determined according to the operation intention, and system operation can be performed on the human-machine interface device according to the system operation instruction, such as controlling the human-machine interface to jump to a specified interface, adjusting parameter settings, controlling industrial equipment, and the like.

[0104] In a feasible implementation, determining a system operation instruction corresponding to the user gesture information can be: identifying a target gesture type according to the gesture information; and generating a system operation instruction corresponding to the target gesture type according to a mapping relationship between a preset gesture type and an operation instruction, to realize device operation.

[0105] It should be noted that the manner of identifying a target gesture type according to gesture information can be: analyzing gesture information (hand image, continuous multiple frames of hand images, or hand motion video) according to a preset gesture recognition model, extracting gesture features, and identifying a target gesture type according to the gesture features.

[0106] Specifically, the gesture features include hand position and motion trajectory (such as three-dimensional coordinates of fingers, palm, and arm and their motion trajectories), hand posture (such as finger bending degree, palm unfolding or fist state, and the like), and the like; and the target gesture type includes raising a thumb, five fingers open, five fingers pinching, and the like.

[0107] Further, the system operation instruction corresponding to the target gesture type is determined according to a mapping relationship between a preset gesture type and an operation instruction, and device operation is realized through the system operation instruction; the mapping relationship can be preset, for example, fist and emergency stop mapping, open palm and start mapping, finger upward sliding and vertical up parameter +1 mapping, OK gesture and confirm / execute command mapping, hand shaking and cancel / exit command mapping, and the like.

[0108] In the above manner, the device can also be operated remotely when the user wears gloves or is inconvenient to approach the human-machine interface device.

[0109] It should be noted that the user image information also includes user face information, and the implementation in which the control instruction is generated according to the interaction information to enable the human-machine interface device to perform device operation can be:

[0110] Compare the user image information with the preset target image information in terms of image similarity; when the image similarity is greater than or equal to the target image similarity, compare the user fingerprint information with the preset target fingerprint information in terms of fingerprint similarity; when the fingerprint similarity is greater than or equal to the target fingerprint similarity, generate a user permission verification instruction; and verify the operation permission of the user according to the user permission verification instruction.

[0111] The embodiment adopts the dual identity verification mechanism of face recognition and fingerprint recognition, and significantly improves the security level of device authorization.

[0112] Specifically, the user image information can be compared with the preset target image information, and when the image similarity is greater than or equal to the target image similarity, it is considered that the portrait feature comparison is passed, and the fingerprint feature is further identified according to the fingerprint information. The fingerprint similarity between the user fingerprint information and the preset target fingerprint information is compared, and when the fingerprint similarity is greater than or equal to the target fingerprint similarity, it is considered that the fingerprint comparison is passed, that is, the current verification user exists in the feature library, but the operation permission corresponding to the user in the feature library is different. Therefore, a user permission verification instruction is further generated, and the operation permission of the user is verified according to the user permission verification instruction, that is, it is judged whether the current user exists operation permission; if any one of the verifications fails, the operation is prohibited and the abnormal event is recorded.

[0113] The target image similarity and the target fingerprint similarity can be 80%, so as to ensure security while allowing certain environmental interference and user operation differences.

[0114] In special scenarios, the permission verification can also be performed according to the fingerprint information and the facial information and their respective weight proportions. For example, when the user cannot touch the fingerprint collection device due to wearing gloves, facial recognition can be used as a supplement. Or in strong light / dark light environment, fingerprint recognition can be used as a supplement. That is, the weight proportions of the fingerprint information and the facial information can be adjusted according to the environmental changes (environmental data is obtained through a sensor to determine the environmental changes), and the permission verification can be flexibly realized.

[0115] It should be noted that, in response to the interactive operation of the user, the implementation manner of obtaining the interactive information according to the interactive operation can be:

[0116] In response to a conference request initiated by a remote device, a network connection with the remote device is established; remote guidance information is obtained according to the conference request, wherein the remote guidance information includes audio and video information; a remote conference operation instruction is generated according to the audio and video information; and a remote conference mode with the remote device is established according to the remote conference operation instruction.

[0117] The man-machine interface device in the embodiment can listen to a conference request from a remote device (such as an expert terminal, a PC, a tablet, etc.) through a wireless network module (such as Wi-Fi, 4G / 5G), the conference request including initiator identity information, conference type (voice / video), encryption identifier or permission verification information, etc., and automatically selecting an optimal communication protocol (such as SIP, WebRTC, RTMP) to establish a connection according to the conference request content, so as to obtain remote guidance information according to the conference request, the remote guidance information including an audio stream (such as remote expert voice), a video stream (such as a remote picture), or a control signal (such as a remote operation suggestion, a marking instruction), etc.

[0118] Further, the remote conference operation instruction is generated according to the audio and video information, and the man-machine interface device can enter a remote conference mode with the remote device according to the remote conference operation instruction: enabling a camera / microphone to collect local images and sounds, receiving and displaying a remote video picture, starting voice recognition, gesture recognition and other auxiliary functions, etc. Thus, efficient cooperation between a remote expert and on-site personnel is realized, and operation efficiency is improved.

[0119] In the embodiment, the device can also be operated remotely when the user wears gloves or is inconvenient to approach the man-machine interface device. The dual identity verification mechanism of "face recognition and fingerprint recognition" significantly improves the security level of device authorization. Remote guidance can be obtained without leaving the site, reducing the round trip time and improving the speed of troubleshooting and device debugging; the safety control of the on-site environment is maintained, and potential safety hazards caused by external devices are avoided.

[0120] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the control method of the multi-modal man-machine interface device of the present application. More forms of simple transformation based on this technical concept are within the protection scope of the present application.

[0121] The present application provides an electronic device, which comprises at least one processor and a memory connected with the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the control method of the multi-modal man-machine interface device in the above-mentioned embodiment one.

[0122] The following refers to Figure 3The diagram illustrates a structural schematic of an electronic device suitable for implementing the embodiments of this application. The electronic devices in the embodiments of this application may include, but are not limited to, mobile terminals such as mobile phones, tablets, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital televisions and desktop computers. Figure 3 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0123] like Figure 3 As shown, the electronic device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the electronic device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. While electronic devices with various systems are shown in the figures, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.

[0124] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer readable medium, the computer program containing program code for executing the method shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiments disclosed in the present application are executed.

[0125] The electronic device provided by the present application adopts the control method of the multi-modal human-computer interface device in the above-mentioned embodiments, and can solve the technical problem of low device operation efficiency. Compared with the prior art, the electronic device provided by the present application has the same beneficial effects as the control method of the multi-modal human-computer interface device provided by the above-mentioned embodiments, and other technical features in the electronic device are the same as the features disclosed in the previous embodiment method, which will not be repeated here.

[0126] It should be understood that parts of the present application can be realized by hardware, software, firmware or a combination thereof. In the description of the above-mentioned embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0127] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0128] The present application provides a computer readable storage medium having stored thereon computer readable program instructions (i.e. computer program) for executing the control method of the multi-modal human-computer interface device in the above-mentioned embodiments.

[0129] The computer readable storage medium provided in the present application may, for example, be a U disk, but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, system, or device, or any combination of the above. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more conductive wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present embodiment, the computer readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer readable storage medium can be transmitted in any suitable medium, including but not limited to electrical wires, optical cables, RF (Radio Frequency), and the like, or any suitable combination of the above.

[0130] The above computer readable storage medium can be contained in an electronic device, or can exist separately without being assembled into an electronic device.

[0131] The above computer readable storage medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the above control method of the multi-modal human-computer interface device.

[0132] Computer program code for carrying out operations of the present application can be written in one or more programming languages or combinations of languages including object oriented programming languages such as Java, Smalltalk, C++ or conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0133] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0134] The modules involved in the embodiments of the present application can be implemented in software or hardware. In some cases, the name of the module does not constitute a limitation on the module itself.

[0135] The computer readable storage medium provided by the present application is a computer readable storage medium, which stores computer readable program instructions (i.e. computer program) for executing the control method of the multi-modal human-computer interface device, and can solve the technical problem of low device operation efficiency. Compared with the prior art, the computer readable storage medium provided by the present application has the same beneficial effects as the control method of the multi-modal human-computer interface device provided by the above embodiments, and will not be described here.

[0136] The present application also provides a computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the control method of the multi-modal human-computer interface device as described above.

[0137] The computer program product provided by the present application can solve the technical problem of low device operation efficiency. Compared with the prior art, the computer program product provided by the present application has the same beneficial effects as the control method of the multi-modal human-computer interface device provided by the above embodiments, and will not be described here.

[0138] The above merely describes some embodiments of the present application, and does not limit the patent scope of the present application, and any equivalent structural transformation, direct / indirect application in other related technical fields made by the present application specification and the drawings contents are included in the patent protection scope of the present application. All the actions of obtaining signals, information or data in the present application are carried out under the premise of complying with the corresponding data protection regulations and policies of the country, and obtaining the authorization given by the corresponding device owner.

Claims

1. A multimodal human-machine interface device, characterized in that, The multimodal human-machine interface device includes: An information acquisition device is used to respond to a user's interactive operation and acquire interactive information based on the interactive operation; the information acquisition device includes a camera module, a fingerprint module, and a remote guidance module; the camera module is used to acquire user image information, the fingerprint module is used to acquire user fingerprint information, and the remote guidance module is used to acquire remote guidance information; the interactive information includes at least one of user image information, fingerprint information, and remote guidance information. A processor device is configured to generate control instructions based on the interaction information, so that the human-machine interface device can operate the device according to the control instructions.

2. The multimodal human-machine interface device as described in claim 1, characterized in that, The user image information includes user gesture information, and the control instructions include system operation instructions; the camera module is used to acquire the user gesture information; the processor device is used to generate system operation instructions based on the user gesture information, so that the human-machine interface device can perform device system operations according to the system operation instructions.

3. The multimodal human-machine interface device as described in claim 2, characterized in that, The user image information also includes user facial information; the processor device is used to generate a user permission verification instruction based on the fingerprint information and the facial information, so that the human-machine interface device can verify the user's operation permission based on the user permission verification instruction.

4. The multimodal human-machine interface device as described in claim 1, characterized in that, The remote guidance module includes: A wireless network module, which is used to establish a network connection with a remote device; Audio and video module, the audio and video module being used to acquire audio and video signals; The processor device is used to respond to a conference request initiated by the remote device and generate remote conference operation instructions based on the audio and video signals, so as to enable the human-machine interface device to establish a remote conference mode with the remote device.

5. A control method for a multimodal human-machine interface device, characterized in that, The control method for the multimodal human-machine interface device as described in any one of claims 1-4 includes: In response to a user's interactive operation, interactive information is obtained based on the interactive operation, wherein the interactive information includes at least one of user image information, fingerprint information, and remote guidance information; Control commands are generated based on the interactive information, so that the human-machine interface device can operate the device according to the control commands.

6. The control method for a multimodal human-machine interface device as described in claim 5, characterized in that, The user image information includes user gesture information, the control instructions include system operation instructions, and the step of generating control instructions based on the interaction information so that the human-machine interface device can operate the device according to the control instructions includes: Determine the system operation command corresponding to the user gesture information; Perform system operations on the human-machine interface device according to the system operation instructions.

7. The control method for a multimodal human-machine interface device as described in claim 6, characterized in that, The user image information also includes user facial information, and the step of generating control commands based on the interaction information to enable the human-machine interface device to operate the device according to the control commands includes: Compare the image similarity between the user image information and the preset target image information; When the image similarity is greater than or equal to the target image similarity, the fingerprint similarity between the user's fingerprint information and the preset target fingerprint information is compared. When the fingerprint similarity is greater than or equal to the target fingerprint similarity, the user permission verification instruction is generated; Verify the user's operation permissions according to the user permission verification instruction.

8. The control method for a multimodal human-machine interface device as described in claim 5, characterized in that, The step of responding to a user's interactive operation and obtaining interactive information based on the interactive operation includes: In response to a conference request initiated by a remote device, establish a network connection with the remote device; The remote guidance information is obtained according to the meeting request, wherein the remote guidance information includes audio and video information; Generate remote conference operation instructions based on the audio and video information; A remote conference mode is established with the remote device according to the remote conference operation instructions.

9. An electronic device, characterized in that, The electronic device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the control method for the multimodal human-machine interface device as claimed in any one of claims 5 to 7.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the control method for the multimodal human-machine interface device as described in any one of claims 5 to 7.

11. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the control method for a multimodal human-machine interface device as described in any one of claims 5 to 7.

Citation Information

Patent Citations

  • Interaction method and device based on multiple modes, storage medium and intelligent screen equipment

    CN111966212A

  • Artificial intelligence interaction method and device, display equipment, storage medium and product

    CN119743637A

  • An audio and video transmission and intelligent interactive management platform based on the Internet of Things

    CN119766851A

  • Method, apparatus, electronic device, and storage for medium extended reality-based interaction control

    EP4509962A1

  • Smart ring providing multi-mode control in a personal area network

    US20190155385A1