Multi-modal interaction method, system, and computer readable medium for unmanned device

CN115937951BActive Publication Date: 2026-09-29EAST CHINA UNIV OF SCI & TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211634892.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-19
Publication Date
2026-09-29
Estimated Expiration
2042-12-19

AI Technical Summary

Technical Problem

[0004]本发明要解决的技术问题是提供一种无人设备的多模态交互方法、系统及计算机可读介质,解决现有交互方法不够智能化以及交互信息融合冗余较大的问题

Benefits of technology

[0020]本发明提出一种基于行为树的多模态交互方法和系统,利用行为树的触发机制开始进行遍历,并根据不同的节点类型执行不同的遍历方式,根据实际的应用需求在相应的条件节点上部署不同模态的交互指令,当遍历到条件节点时,会根据当前模态的识别结果,驱动动作节点对应的无人设备进行相应的动作。本发明可以有效地解决多模态交互过程中交互策略不灵活、交互方法不够智能化以及交互信息融合冗余较大的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115937951B_ABST
    Figure CN115937951B_ABST
Patent Text Reader

Abstract

The application provides a multi-modal interaction method and system of unmanned equipment and a computer readable medium, wherein the method comprises the following steps: acquiring a behavior tree corresponding to a current interaction task of the unmanned equipment, the behavior tree comprising a plurality of condition nodes, and different modal interaction instructions being deployed on each condition node; taking a first control node of the behavior tree as a root node, and starting to traverse the behavior tree from the root node; when one of the condition nodes is reached, identifying an interaction instruction corresponding to the condition node, and driving the unmanned equipment corresponding to the action node to perform a corresponding action according to the identification result. The application can effectively solve the problems of inflexible interaction strategy, insufficient intelligent interaction method and large redundant interaction information fusion in the multi-modal interaction process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates primarily to the field of human-computer interaction, and more particularly to a multimodal interaction method, system, and computer-readable medium for unmanned equipment. Background Technology

[0002] As interactive technologies mature, more and more modal interaction methods are being applied to electronic devices.

[0003] CN 115062131A discloses a multimodal human-computer interaction method and apparatus. This method collects multimodal information input by the user at the terminal, including video information, voice information, and touch screen event information. It then fuses and processes the user's input commands and natural language input, outputting a response. While this method improves human-computer interaction performance by fusing multimodal information, it places high demands on the performance of the natural language processing unit (NLP) and the natural language processing unit (NLP) for scenarios with complex input commands and natural language input. Furthermore, it suffers from insufficient intelligence in the interaction method and significant redundancy in the fusion of interaction information. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a multimodal interaction method, system and computer-readable medium for unmanned equipment, so as to solve the problems of insufficient intelligence and large redundancy in the fusion of interaction information in existing interaction methods.

[0005] To address the aforementioned technical problems, this invention provides a multimodal interaction method for unmanned devices, comprising: acquiring a behavior tree corresponding to the current interaction task of the unmanned device, the behavior tree including multiple condition nodes, each condition node deploying different modal interaction instructions; traversing the behavior tree starting from the first control node of the behavior tree as the root node; when traversing to one of the condition nodes, identifying the interaction instruction corresponding to the condition node, and driving the unmanned device corresponding to the action node to perform a corresponding action based on the identification result.

[0006] Optionally, the interaction instructions include gesture interaction instructions, brain-computer interaction instructions, or lip-reading interaction instructions.

[0007] Optionally, when the modal interaction instruction corresponding to the condition node is a gesture interaction instruction, the step of recognizing the gesture interaction instruction includes: designing multiple gesture actions, each gesture action corresponding to a first control instruction of an unmanned device; acquiring IMU data of hand actions; inputting the IMU data into a gesture detection model to determine whether the hand action is a gesture action; if so, further determining the specific gesture action category through a gesture recognition model.

[0008] Optionally, inputting the IMU data into the gesture detection model includes: performing online parsing and conversion of the IMU data to obtain IMU decimal data; extracting valid active segment data from the IMU decimal data; extracting features from the active segment data; and inputting the extracted features into the gesture detection model.

[0009] Optionally, driving the unmanned device corresponding to the action node to perform the corresponding action based on the recognition result includes: mapping the gesture action category to a first control command of the unmanned device, and sending the first control command to the unmanned device, wherein the first control command belongs to a series of large-amplitude, low-precision motion commands.

[0010] Optionally, when the modal interaction instruction corresponding to the condition node is a brain-computer interaction instruction, the step of identifying the brain-computer interaction instruction includes: designing multiple brainwave signal evoked paradigms, each brainwave signal evoked paradigm corresponding to a second control instruction of an unmanned device; real-time acquisition of brainwave signal data triggered by the brainwave signal evoked paradigm; filtering and downsampling the brainwave signal data to obtain first brainwave signal data; and classifying the first brainwave signal data using a canonical correlation analysis algorithm to obtain the brainwave signal evoked paradigm category.

[0011] Optionally, the step of classifying the first EEG signal data using canonical correlation analysis includes: constructing a corresponding EEG template signal according to the EEG signal evoked paradigm; calculating the correlation coefficient between the first EEG signal data and the EEG template signal using the canonical correlation analysis algorithm, wherein the EEG template signal corresponding to the maximum correlation coefficient is the target signal, and obtaining the EEG signal evoked paradigm category based on the target signal.

[0012] Optionally, driving the unmanned device corresponding to the action node to perform the corresponding action based on the recognition result includes: mapping the EEG signal evoked paradigm category into a second control command for the unmanned device, and sending the second control command to the unmanned device, wherein the second control command belongs to the series of smooth and rapid movement commands.

[0013] Optionally, when the modal interaction instruction corresponding to the condition node is a lip-reading interaction instruction, the steps for recognizing the lip-reading interaction instruction include: designing multiple text instructions, each text instruction corresponding to a third control instruction of an unmanned device; real-time acquisition of the user's lip image sequence data; performing frame segmentation and filtering on the lip image sequence data to obtain first lip image sequence data; extracting features from the first lip image sequence data, and using a softmax classifier to classify the extracted features to obtain the text instruction category.

[0014] Optionally, driving the unmanned device corresponding to the action node to perform the corresponding action based on the recognition result includes: mapping the text instruction category to a third control instruction of the unmanned device, and sending the third control instruction to the unmanned device, wherein the third control instruction belongs to the precise low-speed motion instruction series.

[0015] To address the aforementioned technical problems, this invention provides a multimodal interaction system for unmanned devices, comprising: an unmanned device; an interaction management module configured to acquire a behavior tree corresponding to the current interaction task of the unmanned device, the behavior tree including multiple condition nodes, each condition node deploying different modal interaction instructions, the behavior tree being traversed starting from the first control node of the behavior tree as the root node, and when one of the condition nodes is encountered, a judgment request being sent to the multimodal interaction device; the multimodal interaction device configured to, upon receiving the judgment request, identify the interaction instruction corresponding to the condition node, and drive the unmanned device corresponding to the action node to perform a corresponding action based on the identification result.

[0016] Optionally, the interaction management module includes: a client unit configured to receive a function call sent by a condition node and publish a target task to a service unit when the behavior tree traverses to one of the condition nodes; and a service unit configured to receive the target task, send a judgment request to the multimodal interaction device, receive control instructions fed back by the multimodal interaction device, and send the control instructions to the unmanned device.

[0017] Optionally, the multimodal interaction device includes an IMU data glove, an EEG signal acquisition device, or an augmented reality helmet.

[0018] To address the aforementioned technical problems, the present invention provides a computer-readable medium storing computer program code, which, when executed by a processor, implements the method described above.

[0019] Compared with the prior art, the present invention has the following advantages:

[0020] This invention proposes a multimodal interaction method and system based on behavior trees. It utilizes the triggering mechanism of behavior trees to initiate traversal, executing different traversal methods according to different node types. Based on actual application requirements, different modal interaction commands are deployed on corresponding condition nodes. When a condition node is reached, the corresponding unmanned device is driven to perform an action based on the recognition result of the current modality. This invention effectively solves the problems of inflexible interaction strategies, insufficiently intelligent interaction methods, and excessive redundancy in interaction information fusion during multimodal interaction. Attached Figure Description

[0021] The accompanying drawings are included to provide a further understanding of this application; they are incorporated into and constitute a part of this application. The drawings illustrate embodiments of this application and, together with this specification, serve to explain the principles of the invention. In the drawings:

[0022] Figure 1 This is a flowchart of a multimodal interaction method for unmanned equipment according to an embodiment of the present invention;

[0023] Figure 2 This is a schematic diagram of the behavior tree corresponding to an embodiment of the present invention when there is no tracking target;

[0024] Figure 3 This is a flowchart of a gesture interaction method according to an embodiment of the present invention;

[0025] Figure 4 This is a flowchart of a brainwave interaction method according to an embodiment of the present invention;

[0026] Figure 5 This is a flowchart of a lip-reading interaction method according to an embodiment of the present invention;

[0027] Figure 6 This is a schematic diagram of the behavior tree corresponding to a tracking target according to an embodiment of the present invention;

[0028] Figure 7 This is a system block diagram of a multimodal interaction system for unmanned equipment according to an embodiment of the present invention. Detailed Implementation

[0029] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are merely some examples or embodiments of this application. For those skilled in the art, these drawings can be applied to other similar scenarios without creative effort. Unless obvious from the context or otherwise specified, the same reference numerals in the drawings represent the same structures or operations.

[0030] As indicated in this application and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" are not specifically singular and may include plural forms. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of explicitly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.

[0031] Furthermore, it should be noted that the use of terms such as "first" and "second" to define components is merely for the purpose of distinguishing the corresponding components. Unless otherwise stated, these terms have no special meaning and therefore should not be construed as limiting the scope of protection of this application. In addition, although the terminology used in this application is selected from commonly known and used terms, some terms mentioned in this application's specification may have been chosen by the applicant according to his or her judgment, and their detailed meanings are explained in the relevant sections of this description. Moreover, this application should be understood not only through the actual terms used, but also through the meaning implied by each term.

[0032] Flowcharts are used in this application to illustrate the operations performed by the system according to embodiments of this application. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, various steps can be processed in reverse order or simultaneously. Furthermore, other operations may be added to these processes, or one or more steps may be removed from these processes.

[0033] Figure 1 This is a flowchart of a multimodal interaction method for unmanned equipment according to an embodiment of the present invention. Figure 1 As shown, the multimodal interaction method 100 for unmanned devices includes the following steps:

[0034] Step S11: Obtain the behavior tree corresponding to the current interaction task of the unmanned device. The behavior tree includes multiple condition nodes, and different modal interaction instructions are deployed on each condition node.

[0035] The interaction tasks of unmanned vehicles are divided into two categories based on whether there is a tracking target. When there is no tracking target, the unmanned vehicle is in patrol mode, and its interaction task is to detect the target. When there is a tracking target, the unmanned vehicle is in tracking mode, and its interaction task is to track the target. Figure 2 This is a schematic diagram of the behavior tree corresponding to an embodiment of the present invention when there is no tracked target. For example... Figure 2As shown, the behavior tree 200 corresponding to the absence of a tracking target includes a root node 2 and child nodes. The child nodes include sequence node 21, condition node 211, selection node 212, sequence node A, sequence node B, sequence node C, condition node A1, action node A2, condition node B1, action node B2, condition node C1, and action node C2. Different modalities of interaction commands are deployed on each condition node. Gesture interaction commands are deployed on condition node A1, and action node A2 is the interaction action executed when condition node A1 is successfully judged. The interaction action includes controlling the unmanned device to perform large-amplitude, low-precision movements in the up, down, left, right, forward, and backward directions. EEG interaction commands are deployed on condition node B1, and action node B2 is the interaction action executed when condition node B1 is successfully judged. The interaction action includes controlling the unmanned device to perform smooth, rapid movements in the up, down, left, right, forward, and backward directions. Lip reading interaction commands are deployed on condition node C1, and action node C2 is the interaction action executed when condition node C1 judges that the condition node C1 is successful. The interaction action includes controlling the unmanned device to perform corresponding precise low-speed movements up, down, left, right, forward, and backward.

[0036] Step S12: Traverse the behavior tree starting from the root node.

[0037] Based on the behavior tree triggering mechanism, starting from the root node of the behavior tree, all child nodes are traversed continuously. Different types of child nodes send different function calls, and the child nodes report their current state to the root node through callback functions. When the child node is a selection node, if the child node's callback function returns a success status, all subsequent nodes are paused, and the current child node is set to an idle state; if the current child node is the last node, failure is returned directly. When the child node is a sequential node, if the child node's callback function returns a failure status, all subsequent nodes are paused; otherwise, the current child node is set to an idle state; if the current child node is the last node, success is returned directly.

[0038] by Figure 2To explain, the behavior tree 200 is traversed starting from the root node 2. When the sequential node 21 is reached, according to the definition of a sequential node, the success of the condition node 211 is checked first; in other words, it is checked whether the unmanned device has no tracking target. If so, the condition node 211 returns to a success or running state, and then the selection node 212 is triggered in a left-to-right order. The selection node 212 triggers the lower subtrees in a left-to-right order. The first subtree is the sequential node A, which is determined by the lower condition node A1 to see if there is a gesture interaction command. When there is a corresponding gesture interaction command, it returns to a success or running state and executes the corresponding action node A2 in sequence, controlling the unmanned device to perform large-amplitude, low-precision movements such as up, down, left, right, forward, and backward. The second subtree is the sequential node B, which is determined by the lower condition node B1 to see if there is a brainwave interaction command. When there is a brainwave interaction command, it returns to a success or running state and executes the corresponding action node B2 in sequence, controlling the unmanned device to perform smooth and rapid movements such as up, down, left, right, forward, and backward. The third subtree is the sequential node C. The lower-level condition node C1 determines whether there is a lip-reading interaction command. When there is a lip-reading interaction command, it returns a success status or a running status and sequentially executes the corresponding action node C2 to control the unmanned equipment to perform the corresponding precise low-speed movements of up, down, left, right, forward, and backward.

[0039] Step S13: When traversing to one of the condition nodes, identify the modal interaction command corresponding to the condition node, and drive the unmanned device corresponding to the action node to perform the corresponding action based on the identification result.

[0040] When condition node A1 is reached, it is determined whether there is a gesture interaction command. If there is a corresponding gesture interaction command, a success status or a running status is returned, and the corresponding action node A2 is executed sequentially. Action node A2 includes recognizing the gesture interaction command and driving the corresponding unmanned device to perform the corresponding action based on the recognition result. Figure 3 This is a flowchart of a gesture interaction method according to an embodiment of the present invention. Figure 3 As shown, the gesture interaction method 300 includes the following steps:

[0041] Step S31: Design a variety of gesture actions, each gesture action corresponding to a first control command for an unmanned device.

[0042] Based on the gesture interaction requirements, corresponding gesture actions are pre-designed. For example, ten gesture actions are designed. Users undergo a certain period of training before performing each gesture action. Each gesture action corresponds to a control command. Users perform the corresponding gesture actions by wearing Inertial Measurement Unit (IMU) data gloves according to the task requirements.

[0043] Step S32: Acquire IMU data of hand movements.

[0044] An IMU data glove is used to acquire IMU binary data corresponding to hand movements in real time. In some embodiments, a multi-threaded communication method is used to acquire and read IMU data. In a sub-thread, the IMU data stream is updated in real time in the form of a sliding window. When a gesture is detected, the window data is continuously updated in the main thread for online gesture detection and recognition.

[0045] Step S33: Perform online parsing and conversion of the IMU data to obtain IMU decimal data.

[0046] Specifically, the acquired IMU binary data is parsed to remove the corresponding frame header, frame number, and frame tail, extract the valid IMU binary data, and convert it into decimal data.

[0047] Step S34: Extract valid active segment data from the IMU decimal data.

[0048] The envelope method can be used for active segment detection. The acceleration values ​​of the parsed IMU decimal data have a small range of variation, while the angular velocity values ​​have a large range of variation. Therefore, upper and lower envelopes are drawn for the IMU data of all angular velocity channels. The start and end points of the active segment detection are determined based on the corresponding envelope values.

[0049] Step S35: Extract features from the active segment data and input the extracted features into the gesture detection model.

[0050] The variance of the active segment data can be extracted as a data feature, and the variance of the IMU data of each channel can be calculated to reduce the dimensionality of the data and reflect the characteristics of the gesture.

[0051] Step S36: Determine whether the hand movement is a gesture based on the characteristics. If so, proceed to step S37.

[0052] Features are input into a gesture detection model, which then determines whether a hand movement is a gesture. For example, a Gaussian Bayes model can be used as the gesture detection model, defining ten dynamic gestures as the "gesture" class and one static gesture as the "non-gesture" class, and classifying them based on features.

[0053] Step S37: Further determine the specific gesture category using the gesture recognition model.

[0054] The gesture recognition model can be a random forest model, which can be used to further determine the specific gesture category among the ten gesture categories.

[0055] Step S38: Map the gesture action category to the first control command of the unmanned device, and send the first control command to the unmanned device.

[0056] The first control command belongs to the series of large-amplitude, low-precision motion commands. This series includes large-amplitude, low-precision motion commands for upward, downward, leftward, rightward, forward, and backward movements.

[0057] Continue to refer to Figure 2 When traversing to condition node B1, it checks whether there is a brainwave interaction command. If there is a corresponding brainwave interaction command, it returns a success or running status and sequentially executes the corresponding action node B2. Action node B2 includes recognizing the brainwave interaction command and driving the corresponding unmanned device to perform the corresponding action based on the recognition result. Figure 4 This is a flowchart of a brainwave interaction method according to an embodiment of the present invention. Figure 4 As shown, the brainwave interaction method 400 includes the following steps:

[0058] Step S41: Design multiple EEG signal evoked paradigms, each EEG signal evoked paradigm corresponding to a second control command for an unmanned device.

[0059] Electroencephalogram (EEG) signals can be represented by steady-state visual evoked potentials (SSVEPs). Taking SSVEP signals as an example, various SSVEP signal induction paradigms are pre-designed according to the requirements of brain-computer interfaces based on SSVEP signals. In some embodiments, a paradigm interface consisting of 10 stimulation squares flashing at different frequencies and evenly distributed across the entire screen is designed to induce specific SSVEP signals. Each stimulation square corresponds to a second control command, and the user gazes at the corresponding stimulation square according to the task requirements.

[0060] Step S42: Real-time acquisition of EEG signal data.

[0061] SSVEP signal data evoked by an EEG signal acquisition device is collected in real time using an EEG signal evoked paradigm. In some embodiments, a multi-process communication method is used to acquire and read SSVEP signal data. In the main process, a paradigm program runs, stimulating a cube to flash at different frequencies; in a child process, a backend data acquisition program runs, acquiring and saving the user's SSVEP signal data in real time when the stimulation cube begins to flash, and using the saved data for SSVEP signal instruction recognition when the flashing ends.

[0062] Step S43: Filter and downsample the EEG signal data to obtain the first EEG signal data.

[0063] The EEG signal data of the stimulation square flashing phase is extracted from the SSVEP signal data saved in step S42. The extracted SSVEP signal data is preprocessed by using a Butterworth filter to perform a 4-90Hz bandpass filter on the SSVEP signal of each channel. Then, the filtered data is downsampled to reduce computational complexity.

[0064] Step S44: Classify the first EEG signal data using the canonical correlation analysis algorithm to obtain the EEG signal evoked paradigm category.

[0065] In some embodiments, the step of classifying the first EEG signal data using the Canonical Correlation Analysis (CCA) algorithm includes: constructing corresponding EEG template signals according to the EEG signal evoked paradigm; calculating the correlation coefficient between the first EEG signal data and the EEG template signals using the CCA algorithm; the EEG template signal corresponding to the maximum correlation coefficient is the target signal; and obtaining the EEG signal evoked paradigm category based on the target signal. The CCA algorithm is used to classify SSVEP signals without requiring training. Corresponding SSVEP template signals are constructed based on the flashing frequencies of different stimulus blocks. Then, the correlation coefficient between the SSVEP signal to be identified and each template signal is calculated using the CCA algorithm. The frequency of the template signal corresponding to the maximum correlation coefficient is considered the response frequency of the SSVEP signal, i.e., the target stimulus viewed by the user.

[0066] Step S45: Map the EEG signal evoked paradigm category into a second control command for the unmanned device, and send the second control command to the unmanned device. The second control command belongs to the series of steady and rapid movement commands.

[0067] When unmanned aerial vehicles (UAVs) need to fly smoothly and rapidly, the user focuses on the corresponding stimulus blocks according to the task requirements, generating EEG recognition results. These EEG signals are then converted into secondary control commands for the UAV, enabling smooth and rapid control of the device.

[0068] Continue to refer to Figure 2 When traversing to condition node C1, it checks if there is a lip-reading interaction command. If there is a corresponding lip-reading interaction command, it returns a success or running status and sequentially executes the corresponding action node C2. Action node C2 includes recognizing lip-reading interaction commands and driving the corresponding unmanned device to perform the corresponding action based on the recognition result. Figure 5 This is a flowchart of a lip-reading interaction method according to an embodiment of the present invention. Figure 5 As shown, the lip-reading interaction method 500 includes the following steps:

[0069] Step S51: Design a variety of text instructions, each text instruction corresponding to a third control instruction for an unmanned device.

[0070] Based on the requirements of lip-reading interaction using lip images, corresponding lip-reading text commands are pre-designed. For example, 60 different text commands are designed, each corresponding to a third control command for an unmanned device. The user reads the text commands aloud using lip reading according to the task requirements.

[0071] Step S52: Real-time acquisition of the user's lip image sequence data.

[0072] Augmented reality (AR) headsets can be used to collect real-time image sequences of a user's lips. Specifically, the monocular camera mounted on the AR headset can be used to collect image sequences of lips corresponding to different text commands.

[0073] Step S53: Perform frame segmentation and filtering on the lip image sequence data to obtain the first lip image sequence data.

[0074] In some embodiments, a multi-threaded approach is used to acquire lip image sequence data. In the main thread, the lip image sequence corresponding to the current text command is read in real time. In a sub-thread, the lip movement segment video clip is divided into frames at 30 frames per second to extract the valid movement segment image sequence. The extracted movement segment image sequence is processed by converting the original RGB lip image sequence into a grayscale image sequence, using a notch filter to remove power frequency noise from the lip image sequence, and using a bandpass filter to retain the data with the highest amplitude in the 10-400Hz range, thus obtaining the first lip image sequence data.

[0075] Step S54: Extract features from the first lip image sequence data, and use a softmax classifier to classify the extracted features to obtain the text instruction category.

[0076] The model extracts the Melson coefficient as a data feature from the first lip image sequence to reduce data dimensionality. A 3D convolutional neural network is used to extract spatial information from the lip image sequence. A ResNet residual network is then used to reduce the dimensionality of the high-dimensional lip image sequence. Finally, a gated recurrent unit (GRU) is used to extract temporal information from the lip image sequence. A softmax classifier is then used for real-time recognition of the first lip image sequence. Users read corresponding text instructions according to task requirements, and the end-to-end lip-reading model outputs the text instruction category in real time.

[0077] Step S55: Map the text instruction category to the third control instruction of the unmanned device, and send the third control instruction to the unmanned device. The third control instruction belongs to the precise low-speed motion instruction series.

[0078] When unmanned equipment needs to perform precise low-speed movements, the user reads the corresponding text instructions according to the task requirements and generates lip-reading results. These lip-reading results are then converted into third-party control instructions for the unmanned equipment, enabling precise low-speed control.

[0079] Figure 6 This is a schematic diagram of the behavior tree corresponding to a tracking target according to an embodiment of the present invention. For example... Figure 6 As shown, the behavior tree 600 corresponding to a tracking target includes a root node 6 and child nodes. The child nodes include sequence node 61, condition node 611, selection node 612, sequence node D, sequence node E, sequence node F, condition node D1, action node D2, condition node E1, action node E2, condition node F1, and action node F2. Different modalities of interaction commands are deployed on each condition node. Gesture interaction commands are deployed on condition node D1, and action node D2 is the interaction action executed when condition node D1 is successfully judged. This interaction action includes controlling the unmanned device to perform corresponding tracking actions. Brainwave interaction commands are deployed on condition node E1, and action node E2 is the interaction action executed when condition node E1 is successfully judged. This interaction action includes controlling the unmanned device to perform corresponding obstacle avoidance actions. Lip-reading interaction commands are deployed on condition node F1, and action node F2 is the interaction action executed when condition node F1 is successfully judged. This interaction action includes controlling the unmanned device to perform corresponding encirclement and attack actions.

[0080] When a target is being tracked, the unmanned device is in a chase / track state. First, the behavior tree 600 is traversed starting from the root node 6. When the sequence node 61 is reached, according to its definition, the success of the condition node 611 is checked first; in other words, it's determined whether the unmanned device has a target to track. If so, condition node 611 returns to a success or running state, and then selection node 612 is triggered from left to right. Selection node 612 then triggers its lower subtrees from left to right. The first subtree is sequence node D, where the lower condition node D1 checks for a gesture interaction command. If a corresponding gesture interaction command is present, it returns to a success or running state and sequentially executes the corresponding action node D2, controlling the unmanned device to perform the corresponding tracking action. The second subtree is sequence node E, where the lower condition node E1 checks for a brainwave interaction command. If a brainwave interaction command is present, it returns to a success or running state and sequentially executes the corresponding action node E2, controlling the unmanned device to perform the corresponding obstacle avoidance action. The third subtree is the sequential node F, which determines whether there is a lip-reading interaction command by the lower-level condition node F1. When there is a lip-reading interaction command, it returns a success status or a running status and sequentially executes the corresponding action node F2 to control the unmanned equipment to perform the corresponding encirclement and rescue actions.

[0081] This invention proposes a behavior tree-based multimodal interaction method. The behavior tree begins traversal through a trigger mechanism, executing different traversal methods based on different node types. Different modal interaction commands are deployed on corresponding condition nodes according to actual application requirements. When a condition node is reached, the corresponding unmanned device is driven to perform the appropriate action based on the current modality recognition result. This effectively solves the problems of inflexible interaction strategies, insufficiently intelligent interaction methods, and excessive redundancy in interaction information fusion during multimodal interaction.

[0082] Figure 7 This is a system block diagram of a multimodal interaction system for an unmanned device according to an embodiment of the present invention. Figure 7 As shown, the multimodal interaction system 700 for unmanned devices includes: an interaction management module 71, a multimodal interaction module device 72, and an unmanned device 73. The multimodal interaction device 72 includes, but is not limited to, an IMU data glove, an EEG signal acquisition device, or an augmented reality helmet. The unmanned device 73 includes, but is not limited to, drones, robot dogs, unmanned vehicles, and corresponding unmanned swarm systems. The interaction management module 71 can be implemented in software and deployed in a computer or embedded device. The interaction management module 71 includes a client unit 711 and a service unit 712. The client unit 711 and the service unit 712 communicate through the Robot Operating System (ROS) action communication mechanism. The interaction management module 71 is configured to obtain the behavior tree corresponding to the current interaction task of the unmanned device. The behavior tree includes multiple condition nodes, and each condition node deploys different modal interaction instructions. In some embodiments, the interaction instructions include gesture interaction instructions, brain-computer interaction instructions, or lip-reading interaction instructions.

[0083] The interaction management module 71 uses the first control node of the behavior tree as the root node and traverses the behavior tree starting from the root node. When it encounters a condition node, the condition node sends a function call to the client unit 711. Different condition nodes will send different function calls to the client unit 711. The client unit 711 is configured to, upon receiving a function call, use the ROS action communication mechanism to publish a target task to the service unit 712 or request to cancel the current task. The service unit 712 is also configured to receive the target task sent by the client unit 711, send a judgment request to the multimodal interaction device 72, receive the control commands fed back by the multimodal interaction device 72, and send the control commands to the unmanned device 73.

[0084] The multimodal interaction device 72 is configured to, upon receiving a judgment request, identify the interaction command corresponding to the condition node and drive the unmanned device corresponding to the action node to perform the corresponding action based on the identification result. In some embodiments, the multimodal interaction device 72 includes an instruction recognition unit and an escaping unit. The instruction recognition unit identifies the current interaction command and obtains the recognition result. Based on the Transmission Control Protocol (TCP) communication method, the recognition result is sent to the escaping unit. The escaping unit maps the recognition result to obtain the control command of the unmanned device and feeds the control command back to the service unit 712.

[0085] After receiving the control command, the service unit 712 determines the state of the unmanned device 73. If the unmanned device 73 is in an idle or paused state, the service unit 712 calls a function to the unmanned device 73 through a triggering mechanism, driving the unmanned device 73 to execute the execution thread of the action node and complete the action required by the multimodal interaction command.

[0086] When the unmanned device 73 receives an execution thread request from an action node, it executes the corresponding control command in a thread-blocking manner, and parses and encapsulates the control command to conform to the Micro AirVehicle Link (MAVLink) communication protocol. The encapsulated control command is sent to the unmanned device controller via MAVLink communication to drive the corresponding actuator to complete the action required by the control command. Then, the unmanned device 73 feeds back the execution result to the service unit 712, which in turn feeds it back to the client unit 711.

[0087] This application also includes a computer-readable medium storing computer program code that, when executed by a processor, implements the aforementioned multimodal interaction method for unmanned devices.

[0088] When the multimodal interaction method of unmanned equipment is implemented as a computer program, it can also be stored as an article of manufacture in a computer-readable storage medium. For example, computer-readable storage media can include, but are not limited to, magnetic storage devices (e.g., hard disks, floppy disks, magnetic stripes), optical discs (e.g., compact discs (CDs), digital multifunction discs (DVDs)), smart cards, and flash memory devices (e.g., electrically erasable programmable read-only memory (EPROM), cards, sticks, key drives). Furthermore, the various storage media described herein can represent one or more devices and / or other machine-readable media for storing information. The term "machine-readable medium" can include, but is not limited to, wireless channels and various other media (and / or storage media) capable of storing, containing, and / or carrying code and / or instructions and / or data.

[0089] It should be understood that the embodiments described above are merely illustrative. The embodiments described herein may be implemented in hardware, software, firmware, middleware, microcode, or any combination thereof. For hardware implementation, the processor may be implemented within one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, and / or other electronic units designed to perform the functions described herein, or combinations thereof.

[0090] Some aspects of this application can be executed entirely by hardware, entirely by software (including firmware, resident software, microcode, etc.), or by a combination of hardware and software. The aforementioned hardware or software may be referred to as a "data block," "module," "engine," "unit," "component," or "system." The processor may be one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DAPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, or combinations thereof. Furthermore, aspects of this application may manifest as computer products residing in one or more computer-readable media, including computer-readable program code. For example, computer-readable media may include, but are not limited to, magnetic storage devices (e.g., hard disks, floppy disks, magnetic tapes, etc.), optical discs (e.g., compressed CDs, digital multifunction DVDs, etc.), smart cards, and flash memory devices (e.g., cards, sticks, key drives, etc.).

[0091] A computer-readable medium may contain a propagated data signal containing computer program code, for example, on baseband or as part of a carrier wave. This propagated signal may take various forms, including electromagnetic, optical, and so on, or suitable combinations thereof. A computer-readable medium can be any computer-readable medium other than a computer-readable storage medium, which can be connected to an instruction execution system, apparatus, or device to enable communication, propagation, or transmission of a program for use. The program code located on the computer-readable medium can be propagated through any suitable medium, including radio, cable, fiber optic cable, radio frequency signals, or similar media, or any combination of the above media.

[0092] As indicated in this application and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" are not specifically singular and may include plural forms. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of explicitly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.

[0093] Unless otherwise specifically stated, the relative arrangement, numerical expressions, and values ​​of the components and steps described in these embodiments do not limit the scope of this application. It should also be understood that, for ease of description, the dimensions of the various parts shown in the drawings are not drawn to actual scale. Techniques, methods, and devices known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and devices should be considered part of the specification. In all examples shown and discussed herein, any specific values ​​should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values. It should be noted that similar reference numerals and letters in the following drawings denote similar items; therefore, once an item is defined in one drawing, it need not be further discussed in subsequent drawings.

[0094] Furthermore, it should be noted that the use of terms such as "first" and "second" to define components is merely for the purpose of distinguishing the corresponding components. Unless otherwise stated, these terms have no special meaning and therefore should not be construed as limiting the scope of protection of this application. In addition, although the terminology used in this application is selected from commonly known and used terms, some terms mentioned in this application's specification may have been chosen by the applicant according to his or her judgment, and their detailed meanings are explained in the relevant sections of this description. Moreover, this application should be understood not only through the actual terms used, but also through the meaning implied by each term.

[0095] The basic concepts have been described above. Obviously, for those skilled in the art, the above disclosure is merely illustrative and does not constitute a limitation of this application. Although not explicitly stated herein, those skilled in the art may make various modifications, improvements, and corrections to this application. Such modifications, improvements, and corrections are suggested in this application, and therefore remain within the spirit and scope of the exemplary embodiments of this application.

Claims

1. A multimodal interaction method for unmanned equipment, characterized in that, include: Obtain the behavior tree corresponding to the current interaction task of the unmanned device. The behavior tree includes a root node and child nodes. The child nodes include a first sequence node, a first condition node, a selection node, a second sequence node, a second condition node, and an action node. Different modal interaction instructions are deployed on each second condition node. The interaction instructions include gesture interaction instructions, brain-computer interaction instructions, or lip-reading interaction instructions. The behavior tree is traversed starting from the first control node of the behavior tree. When a second condition node is encountered, the interaction command corresponding to the second condition node is identified, and the unmanned device corresponding to the action node is driven to perform the corresponding action based on the identification result. The traversal of the behavior tree includes: starting from the root node of the behavior tree, continuously traversing all child nodes. When traversing to the first sequence node, first determine whether the first condition node is successful. If the first condition node returns a successful state, a selection node will be triggered. The selection node will trigger the second sequence node in order from left to right. The second sequence node is determined by the lower-level second condition node to determine whether there is an interaction instruction. When there is an interaction instruction, the action node is executed to control the unmanned device to perform the corresponding action. Wherein, when the current interaction task is to detect a target, there is a corresponding behavior tree without a tracked target. In the behavior tree without a tracked target, the first condition node is used to determine whether the unmanned device has no tracked target. When the current interaction task is to track a target, there is a corresponding behavior tree with a tracked target. In the behavior tree with a tracked target, the first condition node is used to determine whether the unmanned device has a tracked target. In the target tracking behavior tree, when the second condition node is equipped with a gesture interaction command, the action node controls the unmanned device to perform a corresponding tracking action; when the second condition node is equipped with an EEG interaction command, the action node controls the unmanned device to perform a corresponding obstacle avoidance action; and when the second condition node is equipped with a lip-reading interaction command, the action node controls the unmanned device to perform a corresponding encirclement and reinforcement action.

2. The method as described in claim 1, characterized in that, When the modal interaction command corresponding to the second condition node is a gesture interaction command, the steps for recognizing the gesture interaction command include: Design a variety of hand gestures, each of which corresponds to a first control command for the unmanned device; Acquire IMU data of hand movements; The IMU data is input into the gesture detection model to determine whether the hand movement is a gesture. If so, the gesture recognition model is used to further determine the specific gesture category.

3. The method as described in claim 2, characterized in that, The actions driven by the recognition results to drive the unmanned device corresponding to the action node to perform corresponding actions include: mapping the gesture action category to a first control command of the unmanned device, and sending the first control command to the unmanned device, wherein the first control command belongs to a series of large-amplitude, low-precision motion commands.

4. The method as described in claim 2, characterized in that, The input of the IMU data into the gesture detection model includes: The IMU data is parsed and converted online to obtain IMU decimal data; Extract valid active segment data from the IMU decimal data; Feature extraction is performed on the active segment data, and the extracted features are input into the gesture detection model.

5. The method as described in claim 1, characterized in that, When the modal interaction instruction corresponding to the second condition node is a brain-computer interaction instruction, the steps for recognizing the brain-computer interaction instruction include: Design multiple brain signal evoked paradigms, each of which corresponds to a second control command for an unmanned device; Real-time acquisition of EEG signal data triggered by the aforementioned EEG signal evoked paradigm; The EEG signal data is filtered and downsampled to obtain the first EEG signal data; The first EEG signal data was classified using canonical correlation analysis algorithm to obtain the EEG signal evoked paradigm category.

6. The method as described in claim 5, characterized in that, The steps for classifying the first EEG signal data using canonical correlation analysis include: Construct a corresponding EEG template signal based on the EEG signal evoked paradigm; The correlation coefficient between the first EEG signal data and the EEG template signal is calculated using the canonical correlation analysis algorithm. The EEG template signal corresponding to the maximum correlation coefficient is the target signal. The EEG signal evoked paradigm category is obtained based on the target signal.

7. The method as described in claim 5, characterized in that, The actions driven by the recognition results to drive the unmanned device corresponding to the action node to perform the corresponding actions include: mapping the EEG signal evoked paradigm category into a second control command for the unmanned device, and sending the second control command to the unmanned device. The second control command belongs to the series of smooth and rapid movement commands.

8. The method as described in claim 1, characterized in that, When the modal interaction command corresponding to the second condition node is a lip-reading interaction command, the steps for recognizing the lip-reading interaction command include: Design a variety of text commands, each of which corresponds to a third control command for the unmanned equipment; Real-time acquisition of user lip image sequence data; The lip image sequence data is segmented and filtered to obtain the first lip image sequence data; Feature extraction is performed on the first lip image sequence data, and the extracted features are classified using a softmax classifier to obtain the text instruction category.

9. The method as described in claim 8, characterized in that, The actions driven by the recognition results to drive the unmanned equipment corresponding to the action node to perform corresponding actions include: mapping the text instruction category to the third control instruction of the unmanned equipment, and sending the third control instruction to the unmanned equipment, wherein the third control instruction belongs to the precise low-speed motion instruction series.

10. A multimodal interaction system for unmanned equipment, characterized in that, include: Unmanned equipment; The interaction management module is configured to obtain the behavior tree corresponding to the current interaction task of the unmanned device. The behavior tree includes a root node and child nodes. The child nodes include a first sequence node, a first condition node, a selection node, a second sequence node, a second condition node, and an action node. Different modal interaction instructions are deployed on each second condition node. The interaction instructions include gesture interaction instructions, brain-computer interaction instructions, or lip-reading interaction instructions. The first control node of the behavior tree is taken as the root node. The behavior tree is traversed from the root node. When a second condition node is traversed, a judgment request is sent to the multimodal interaction device. A multimodal interaction device is configured to, upon receiving the judgment request, identify the interaction command corresponding to the second condition node, and drive the unmanned device corresponding to the action node to perform the corresponding action based on the identification result; The traversal of the behavior tree includes: starting from the root node of the behavior tree, continuously traversing all child nodes. When traversing to the first sequence node, first determine whether the first condition node is successful. If the first condition node returns a successful state, a selection node will be triggered. The selection node will trigger the second sequence node in order from left to right. The second sequence node is determined by the lower-level second condition node to determine whether there is an interaction instruction. When there is an interaction instruction, the action node is executed to control the unmanned device to perform the corresponding action. Wherein, when the current interaction task is to detect a target, there is a corresponding behavior tree without a tracked target. In the behavior tree without a tracked target, the first condition node is used to determine whether the unmanned device has no tracked target. When the current interaction task is to track a target, there is a corresponding behavior tree with a tracked target. In the behavior tree with a tracked target, the first condition node is used to determine whether the unmanned device has a tracked target. In the target tracking behavior tree, when the second condition node is equipped with a gesture interaction command, the action node controls the unmanned device to perform a corresponding tracking action; when the second condition node is equipped with an EEG interaction command, the action node controls the unmanned device to perform a corresponding obstacle avoidance action; and when the second condition node is equipped with a lip-reading interaction command, the action node controls the unmanned device to perform a corresponding encirclement and reinforcement action.

11. The system as claimed in claim 10, characterized in that, The interactive management module includes: The client unit is configured to receive a function call sent by one of the second condition nodes and publish the target task to the service unit when the behavior tree traverses to one of the second condition nodes. The service unit is configured to receive the target task, send a judgment request to the multimodal interaction device, receive control instructions fed back by the multimodal interaction device, and send the control instructions to the unmanned device.

12. The system as described in claim 10, characterized in that, The multimodal interaction device includes an IMU data glove, an EEG signal acquisition device, or an augmented reality helmet.

13. A computer-readable medium storing computer program code that, when executed by a processor, implements the method as claimed in any one of claims 1-9.