Multi-mode intention understanding and intelligent navigation method, system and device and storage medium

Through multimodal intention understanding and intelligent navigation methods, combined with voice, gesture and visual information, a multimodal intention understanding algorithm based on HCRF is constructed, which solves the problems of insufficient real experience of students' operation and unnatural interaction in the traditional experimental teaching model, and achieves more efficient experimental teaching effects and personalized learning experience.

CN120197121APending Publication Date: 2025-06-24SHANDONG XIEHE UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510202242.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

In the traditional experimental teaching model, students have insufficient real experience in operation and unnatural interaction process, resulting in limited improvement in experimental teaching effects, making it difficult to meet the needs of modern education for personalized, interactive and immersive learning experience.

Method used

Using multimodal intention understanding and intelligent navigation methods, a hidden sequence extraction algorithm based on multimodal long and short-term memory is constructed by collecting and preprocessing voice, gestures and visual information, and a multimodal intention understanding algorithm based on HCRF is implemented to realize intelligent navigation of user intentions.

Benefits of technology

It improves the real experience of students' operation in the experimental teaching process, improves the interactive process, enhances the effect of experimental teaching, and meets the needs of modern education for personalized, interactive and immersive learning experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197121A_ABST
    Figure CN120197121A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal intention understanding and intelligent navigation method, system and device and a storage medium, and mainly relates to the technical field of multi-modal fusion. Comprising the following steps: collecting multi-modal data information, and preprocessing the collected data information; constructing a hidden sequence extraction algorithm based on multi-modal long and short-term memory for the processed data information; according to the above algorithm, a multi-mode intention understanding algorithm based on HCRF is executed. The system has the beneficial effects that the real operation experience of students in the experiment teaching process is improved, the interaction process is perfected, and the experiment teaching effect is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multimodal fusion, and specifically to a multimodal intention understanding and intelligent navigation method, system, device, and storage medium. Background Art

[0002] The traditional experimental teaching mode has long faced prominent problems such as insufficient real experience for students in operations and unnatural interaction processes, which restricts the improvement of experimental teaching effects.

[0003] Specifically, the existing mode mostly relies on preset experimental steps and fixed instrument operations. Students often passively follow the established procedures, lacking opportunities for independent exploration and actual operation, and it is difficult to truly understand and master experimental principles and methods. In addition, the interaction methods in the traditional experimental environment are single and rigid, with insufficient interaction between students and experimental equipment, teachers, and other students, making it difficult to stimulate learning interest and initiative. This mode not only restricts the development of students' practical abilities and innovative thinking but also fails to meet the needs of modern education for personalized, interactive, and immersive learning experiences.

[0004] Therefore, there is an urgent need for a method of multimodal intention understanding and intelligent navigation based on HCRF+LSTM to solve the above problems. Summary of the Invention

[0005] The purpose of the present invention is to provide a multimodal intention understanding and intelligent navigation method, system, device, and storage medium, which improves the real experience of students in the experimental teaching process, while perfecting the interaction process, and thus enhances the experimental teaching effect.

[0006] To achieve the above object, the present invention is realized through the following technical solutions:

[0007] On the one hand, a multimodal intention understanding and intelligent navigation method is provided, including the following steps:

[0008] S1: Collect multimodal data information and preprocess the collected data information;

[0009] S2: For the processed data information, construct a hidden sequence extraction algorithm based on multimodal long short-term memory;

[0010] S3: According to the algorithm in step S2, execute a multimodal intention understanding algorithm based on HCRF.

[0011] Preferably, in step S1, collecting multimodal data information includes: collecting voice information, collecting gesture data information, and collecting visual information.

[0012] Preferably, in step S1, preprocessing the collected data information includes:

[0013] Use MFCC to extract features from the collected voice information, and the voice information is expressed as: x voice,t = MFCC(voice data);

[0014] Normalize and extract features from the gesture data information to obtain the motion trajectory of the gesture, and the gesture data information is expressed as: x gesture,t = [x 坐标 , x 方向 , x 速度 ;

[0015] Use the fixation duration of the user's fixation point and the fixation frequency within a set time to process the visual information, and the visual information is expressed as: x eye,t = [x 注视时长 , x 注视频率 ;

[0016] Perform a concatenation operation on the above preprocessed multi-modal feature vectors to form a joint observation sequence x t = [x voice,t , x gesture,t , x eye,t , and obtain an observation sequence, expressed as: x = [x1, x2,..., x T .

[0017] Preferably, in step S2, a hidden sequence extraction algorithm based on multi-modal long short-term memory is constructed, including the following steps:

[0018] S21: Use a hidden conditional random field to label the multi-modal operation behavior sequence of the user, and set an observation vector x t for each time period, where t = 1, 2, 3…, all the observation vectors form an observation sequence x = {x1, x2, x3,…x T}, and there is a label sequence y corresponding to it. The elements in the label sequence y represent the operation behaviors of the user. Set a hidden sequence h = {h1, h2, h3,…h T}, and the elements in the hidden sequence represent the relationships between the three modalities of the user;

[0019] S22: Use the input gate of LSTM to control the input degree of new information, input the observation vector x t and the hidden state h t-1 of the previous time step, and determine the acceptance degree of the new information in the input observation vector x t :

[0020] i t = σ(W xi ·x t + W hi ·ht-1 +b i ) (1)

[0022] Use the forget gate of the LSTM to control the retention degree of memory information. Through the input observation vector x t and the hidden state h of the previous time step t-1 , determine the amount of data for retaining memory information:

[0023] f t = σ(W xf ·x t + W hf ·h t-1 + b f ) (2)

[0024] S23: Use the tanh function to determine the storage amount of new information and memory information, and calculate the new cell state through the new observation vector and the hidden state h of the previous time step t-1 :

[0025] c t = f t ⊙ c t-1 + i t ⊙ tanh(W xc ·x t + W hc ·h t-1 + b c ) (3)

[0027] S24: Calculate the hidden state h t by using the output gate o t1 and the tanh activation function:

[0028] h t = o t ⊙ tanh(c t )(4).

[0029] Preferably, in the step S3, execute the multi-modal intention understanding algorithm based on HCRF, including the following steps:

[0030] S31: Define a state transition matrix A and a backtracking table B, where the elements in A are the probabilities of transferring from one hidden state to another, and the calculation method is:

[0031]

[0032] The backtracking table B is used to record, at each time step t, for each possible hidden state h t the optimal hidden state h from the previous time step t-1t-1 ; B(t, h t ) stores the hidden state h that maximizes the state score V(t, h t ) at time step t - 1; t-1 ;

[0033] S32: Initialize the state score table V and the backtracking table B:

[0034] V(1, h1) = π(h1)·f(x1, h1) (6)

[0036] B(1, h1) = None (7)

[0038] For the subsequent time steps t = 2, 3, …, T, continuously update the state score table V and the backtracking table B, and the calculation method is:

[0039]

[0040] S33: Backtrack through the backtracking table B to trace the optimal path, and the calculation method is:

[0041]

[0042] For each time step t = T - 1, T - 2, …, 1, obtain the optimal predecessor hidden state from the backtracking table B, and the calculation method is:

[0043]

[0044] S34: Output the optimal label sequence

[0045] On the other hand, provide a multi-modal intention understanding and intelligent navigation system, including:

[0046] A data acquisition and preprocessing module, used for: acquiring multi-modal data information and preprocessing the acquired data information;

[0047] A first algorithm construction module, used for: constructing a hidden sequence extraction algorithm based on multi-modal long short-term memory for the processed data information;

[0048] A second algorithm construction module, used for: executing a multi-modal intention understanding algorithm based on HCRF.

[0049] On the other hand, provide a multi-modal intention understanding and intelligent navigation device, including a processor and a memory for storing a computer program. When the processor executes the computer program, the steps of the method according to any one of claims 1 - 5 are implemented.

[0050] On the other hand, a computer-readable storage medium has a computer program stored thereon. When the computer program is executed by one or more processors, the steps of the method according to any one of claims 1-5 are implemented.

[0051] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0052] 1. The present invention integrates the user's gesture, voice, and visual information, solving the problems of the single interaction mode and unnatural interaction of the traditional intelligent experimental platform.

[0053] 2. The present invention constructs an intelligent navigation system based on the user's intention, achieving the goal of human-machine collaboration. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 is a flowchart of the method of the present invention;

[0055] Figure 2 is an algorithm framework diagram of the present invention;

[0056] Figure 3 is a schematic diagram of the system structure of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the present application.

[0058] In the present invention, terms such as "upper", "lower", "left", "right", "front", "rear", "vertical", "horizontal", "side", "bottom", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. They are only relational terms determined for the convenience of describing the structural relationship of each component or element of the present invention, and do not specifically refer to any component or element of the present invention. It should not be construed as a limitation of the present invention.

[0059] In the present invention, terms such as "fixed connection", "connected", "connected" should be understood in a broad sense, which can mean a fixed connection, an integral connection or a detachable connection; it can be directly connected or indirectly connected through an intermediate medium. For those skilled in the relevant scientific research or technology in the field, the specific meanings of the above terms in the present invention can be determined according to specific circumstances, and it should not be construed as a limitation of the present invention.

[0060] Embodiment:

[0061] As Figure 1As shown in the figure, this embodiment provides a multi-modal intention understanding and intelligent navigation method, including the following steps:

[0062] S1: Collect multi-modal data information and preprocess the collected data information;

[0063] S2: For the processed data information, construct a hidden sequence extraction algorithm based on multi-modal long short-term memory;

[0064] S3: According to the algorithm in step S2, execute a multi-modal intention understanding algorithm based on HCRF.

[0065] Among them, collecting multi-modal data information in step S1 includes: collecting voice information, collecting gesture data information, and collecting visual information:

[0066] For voice information, in this embodiment, the form of "verb" + "operation object" is adopted. Specifically, keywords are extracted through Baidu Smart Voice Assistant, the voice signal is converted into digital information, and preprocessing is performed;

[0067] For gesture data information, in this embodiment, gesture capture and tracking are performed through a HoloLens2 camera, and the position and movement trajectory of the hand in the gesture information are extracted;

[0068] For visual information, in this embodiment, the gaze duration of the user's gaze point and the gaze frequency of the user's region of interest are obtained through the infrared camera of HoloLens2;

[0069] In addition, in this embodiment, a set of intelligent interaction components are designed based on 3D printing technology, and the appearance of the intelligent interaction components is modeled using UnityProBuilder. The intelligent interaction components have characteristics similar to the shape, size, and grip of real chemical experimental equipment, thereby improving the authenticity of the experiment. And according to different application scenarios, various sensor devices are embedded in the intelligent interaction components, including a position sensor HC-SR04, a nine-axis attitude sensor BMI088, a touch sensor FSR402, and an RFID-RC522. These sensors are distributed at the bottom, inside, and surface of the equipment and can perceive the user's behavior;

[0070] In the experiment, the user's behavior generates four types of data. Taking the sodium-water reaction experiment as an example in this embodiment, one of the operation steps in the experiment is to add water to a beaker. Then, for voice commands such as "add water" and "pour water" issued by the user in this environment, the movement trajectory, direction, and speed of the user's hand in the user gesture data are used as indicators of gesture information. For visual information, the change in the fixation frequency of the region of interest by the user within this period is calculated to reflect the user's visual information. For the processing of voice information, MFCC (Mel Frequency Cepstral Coefficients) is used to extract voice features, and the voice information is represented as x voice,t = MFCC(voice data); The gesture data is standardized and feature-extracted, and the movement trajectory of the gesture is obtained. The gesture information is represented as x gesture,t = [x 坐标 , x 方向 , x 速度 ; For visual information, using the fixation duration of the user's fixation point and the fixation frequency within a specific time, the visual information is represented as x eye,t = [x 注视时长 , x 注视频率 ;

[0071] By splicing the multi-modal feature vectors together, a joint observation sequence x t = [x voice,t , x gesture,t , x eye,t is formed. Therefore, the final observation sequence is represented as x = [x1, x2,..., x T ; For sensor information, the system determines whether the student has operated on the component through the signal change of the sensor, and judges the moving distance and direction of the component through the distance sensor.

[0072] In order to make full use of the long-term dependence of the user's behavior and various modal information, this embodiment uses the method of Hidden Conditional Random Field (HCRF). Hidden Conditional Random Field (HCRF) is an undirected graph that can be used for sequence labeling. Different from ordinary conditional random fields, Hidden Conditional Random Field not only considers the relationship between the observed input sequence and the output label sequence, but also considers the influence of unobserved hidden variables in the model. In addition, the user's behavior also has strong dynamic variability, and traditional static decoding methods are difficult to accurately capture these changes, resulting in a decrease in the accuracy of intention understanding. The dynamic decoding strategy can update the state transition matrix in real time to adapt to the changes in the user's behavior. Therefore, in order to obtain the user's intention more accurately, this embodiment also introduces a dynamic decoding strategy into HCRF to obtain a method based on dynamic decoding HCRF to obtain the user's intention. The framework of the algorithm is as Figure 2 shown;

[0073] Use the method based on dynamic decoding HCRF to label the user's multi-modal operation behavior sequence in the AR experiment. There is an observation vector x at each time period t , t = 1, 2, 3…, all the observation vectors form an observation sequence x = {x1, x2, x3, …x T}, and there is a corresponding label sequence y. The elements in the label sequence represent the operation behaviors of our users. Assume there is a hidden sequence h = {h1, h2, h3, …h T}, and the elements in the hidden sequence represent the relationships between the three modalities of the user;

[0074] In LSTM, the input gate controls the degree of input of new information. Input the observation vector x t and the hidden state h at the previous time step t-1 , and determine how much new information from the input x t should be accepted through formula (1):

[0075] i t = σ(W xi ·x t + W hi ·h t-1 + b i )#(1)

[0076] The forget gate controls the degree of retention of previous memory information. We determine how much previous memory information should be retained by inputting the observation vector x t and the hidden state h at the previous time step t-1 , and through formula (2):

[0077] f t = σ(W xf ·x t + W hf ·h t-1 + b f )#(2)

[0078] Similarly, for the calculation of the output gate, we use the same method, as shown in formula (3):

[0079] o t = σ(W xo ·x t + W ho ·h t-1 + b o )#(3)

[0080] The cell state is the core part of the LSTM network. It is used to transmit long-term dependency information. At this step, first use the input gate i t and the forget gate f tTo update the cell state, the input gate controls the influence of new input information, and the forget gate controls the influence of previous memory information. Then, the tanh function is used to determine how much new information and previous memory information should be stored. Finally, a new cell state is calculated based on the new observation vector and the hidden state h at the previous time step t-1 The new cell state is calculated as shown in Equation (4):

[0081] c t = f t ⊙ c t-1 + i t ⊙ tanh(W xc · x t + W hc · h t-1 + b c ) #(4)

[0082] Finally, the hidden state h is calculated using the output gate o t and the tanh activation function, as shown in Equation (5): t

[0083] h t = o t ⊙ tanh(c t ) #(5)

[0084] Considering the temporal nature of user behavior, for example, user gesture behavior and eye movement data exhibit different characteristics over time. Therefore, the LSTM is used to obtain the hidden state h. The LSTM has memory cells and a gating mechanism, which can effectively handle long-term dependence problems, making it perform well in processing time series data. This concatenation operation of the three modal information can be regarded as feature-level fusion of different modal data, thus providing a unified input format for the subsequent LSTM model. This joint feature vector can more comprehensively represent the comprehensive state of the user at the current time step, enabling the LSTM model to more accurately capture the hidden intentions of the user

[0085] Using the hidden sequence h generated by the LSTM, the dynamic decoding HCRF method is used to further decode the comprehensive state of the user at the current moment, and then accurately infer the user's intention

[0086] First, a state transition matrix A and a backtracking table B are defined. The elements in A represent the probabilities of transitioning from one hidden state to another, and are calculated as shown in Equation (6):

[0087]

[0088] The backtracking table B is used to record, at each time step t, for each possible hidden state h t tThe optimal hidden state h from the previous time step t-1 t-1 , B(t, h t ) stores the hidden state h that maximizes the state score V(t, h t ) at time step t-1; t-1 ;

[0089] Next, initialize the state score table V and the backtracking table B, and the formula is as follows:

[0090] V(1, h1) = π(h1)·f(x1, h1) #(7)

[0091] B(1, h1) = None #(8)

[0092] For the next time steps t = 2, 3,..., T, continuously update the state score table V and the backtracking table B, and the calculation formula is as follows:

[0093]

[0094] Then, backtrack through the backtracking table B to trace the optimal path, and the calculation formula is as follows:

[0095]

[0096] For each time step t = T-1, T-2,..., 1, obtain the optimal precursor hidden state from the backtracking table B, and the calculation formula is as follows:

[0097]

[0098] Finally, output the optimal label sequence

[0099] As Figure 3 shown, this embodiment also provides a multimodal intent understanding and intelligent navigation system, including:

[0100] A data collection and preprocessing module, used for: collecting multimodal data information and preprocessing the collected data information;

[0101] A first algorithm construction module, used for: constructing a hidden sequence extraction algorithm based on multimodal long short-term memory for the processed data information;

[0102] A second algorithm construction module, used for: executing a multimodal intent understanding algorithm based on HCRF.

[0103] This embodiment also provides a device, including:

[0104] At least one processor;

[0105] At least one memory, used for storing at least one program;

[0106] When the at least one program is executed by at least one processor, the at least one processor implements a multimodal intent understanding and intelligent navigation method.

[0107] The content in the above method embodiments is applicable to the device embodiments of the present invention. The functions specifically implemented by the device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0108] In some alternative embodiments, the embodiments presented and described in the steps of the present invention are provided by way of example for the purpose of providing a more comprehensive understanding of the technology. The disclosed method is not limited to the operations and logical flows presented herein. Alternative embodiments are contemplated, in which the order of various operations is changed and sub-operations described as part of a larger operation are performed independently.

[0109] Furthermore, although the present invention has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the functions and / or features described may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It can also be understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present invention. Rather, considering the attributes, functions, and internal relationships of the various functional modules in the device disclosed herein, the actual implementation of the module will be understood within the ordinary skills of an engineer. It can also be understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.

[0110] This embodiment also provides a storage medium storing a program, and the program, when executed by a processor, implements a multimodal intent understanding and intelligent navigation method.

[0111] Similarly, it can be seen that the content in the above method embodiments is applicable to the storage medium embodiments of the present invention, and the functions and beneficial effects achieved are the same as those of the method embodiments.

[0112] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0113] The steps in the embodiments represent or are otherwise described herein as logical and / or steps. For example, they can be considered as a predefined sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in combination with an instruction execution system, apparatus, or device.

[0114] More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection part (electronic device) having one or more wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, a computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other appropriate processing as necessary, and then stored in a computer memory.

[0115] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the described embodiments. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.

Claims

1. A multimodal intent understanding and intelligent navigation method, characterized in that: The following steps are involved: S1: Collect multimodal data information and preprocess the collected data information; S2: For the processed data information, construct a hidden sequence extraction algorithm based on multimodal long short-term memory; S3: According to the algorithm in step S2, execute the HCRF-based multimodal intent understanding algorithm.

2. The multimodal intent understanding and intelligent navigation method according to claim 1, characterized in that: In the step S1, collecting multimodal data information includes: collecting voice information, collecting gesture data information and collecting visual information.

3. The multimodal intention understanding and intelligent navigation method according to claim 2, characterized in that: In step S1, the collected data information is preprocessed, including: MFCC is used to extract features from the collected speech information, and the speech information is represented as: voice,t =MFCC(speech data); The gesture data information is standardized and feature extracted to obtain the motion trajectory of the gesture. The gesture data information is represented as: gesture,t =[x 坐标 ,x 方向 ,x 速度 ]; The visual information is processed using the gaze duration of the user's gaze point and the gaze frequency within a set time. The visual information is represented as: eye,t =[x 注视时长 ,x 注视频率 ]; The preprocessed multimodal feature vectors are concatenated to form a joint observation sequence x t =[x voice,t ,x gesture,t ,x eye,t ], and obtain the observation sequence, expressed as: x = [x1, x2, ..., x T ].

4. The multimodal intention understanding and intelligent navigation method according to claim 1, characterized in that: In step S2, a hidden sequence extraction algorithm based on multimodal long short-term memory is constructed, comprising the following steps: S21: Use hidden conditional random fields to mark the user's multimodal operation behavior sequence and set an observation vector x for each time period t , where t = 1, 2, 3, ..., all observation vectors form an observation sequence x = {x1, x2, x3, ...x T }, and there is a label sequence y corresponding to it, the elements in the label sequence y represent the user's operation behavior, and a hidden sequence h = {h1, h2, h3, ... h T }, the elements in the hidden sequence represent the relationship between the three modalities of the user; S22: Use the input gate of LSTM to control the input degree of new information and input observation vector x t and the hidden state h at the previous time step t-1 , and determine the input observation vector x t Acceptance of new information in: i t =σ(W xi ·x t +W hi ·h t-1 +b i ) (1) The forget gate of LSTM is used to control the retention degree of memory information, by inputting the observation vector x t and the hidden state h of the previous time step t-1 , determine the amount of data to retain memory information: f t =σ(W xf ·x t +W hf ·h t-1 +b f ) (2) S23: Use the tanh function to determine the storage capacity of new information and memory information, and use the new observation vector and the hidden state h of the previous time step t-1 Calculate the new cell state: c t =f t ⊙c t-1 +i t ⊙tanh(W xc ·x t +W hc ·h t-1 +b c ) (3) S24: By using the output gate o t and tanh activation function to calculate the hidden state h t : h t =o t ⊙tanh(c t ) (4)。 5. The multimodal intention understanding and intelligent navigation method according to claim 1, characterized in that: In step S3, executing the HCRF-based multimodal intent understanding algorithm includes the following steps: S31: Define a state transfer matrix A and a lookback table B, where the probability of an element in A transferring from one hidden state to another hidden state is calculated as: The lookback table B is used to record at each time step t, for each possible hidden state h t The optimal hidden state h from the previous time step t-1 t-1 ; B(t,h t ) stores the state score V(t,h at time step t-1, t )The largest hidden state h t-1 ; S32: Initialize the state score table V and the backtracking table B: V(1,h1)=π(h1)·f(x1,h1) (6) B91, h1)=None (7) For the next time steps t=2, 3, ..., T, the state score table V and the backtracking table B are continuously updated, and the calculation method is: S33: Backtrack through the backtracking table B to track the optimal path. The calculation method is: For each time step t = T-1, T-2, ..., 1, the optimal predecessor hidden state is obtained from the lookback table B, and the calculation method is: S34: Output the optimal label sequence 6. Multimodal intention understanding and intelligent navigation system, characterized by: include: The data acquisition and preprocessing module is used to: acquire multimodal data information and preprocess the acquired data information; The first algorithm construction module is used to: construct a hidden sequence extraction algorithm based on multimodal long short-term memory for the processed data information; The second algorithm building module is used to: execute a multimodal intent understanding algorithm based on HCRF.

7. A multimodal intention understanding and intelligent navigation device, characterized in that: The method comprises a processor and a memory for storing a computer program, wherein when the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by one or more processors, the steps of the method according to any one of claims 1 to 5 are implemented.