Experimental intelligent navigation method and system based on multi-modal intention understanding
Through the experimental intelligent navigation method of multimodal intention understanding, combined with the multimodal fusion algorithm of dynamic game and Bayesian network model, the problem of difficulty in realizing natural interaction and human-machine collaboration in existing systems is solved, and intelligent navigation and security guarantees of chemical experiments are realized.
Patent Information
- Application Number
- CN202510218035.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-27
AI Technical Summary
The existing chemical experiment intelligent navigation system is difficult to achieve natural interaction, and is mainly machine-centric, so it is unable to effectively complete the human-machine collaboration experimental process.
An experimental intelligent navigation method based on multimodal intention understanding is adopted, through the fusion of visual mode and sensor mode, combined with a multimodal fusion algorithm of dynamic game, an experimental auxiliary model is constructed to realize the experimental process of human-machine collaboration.
The intelligentization of the experimental process is achieved through the collaborative mode between humans and machines, dynamically adjusting the weight of modal data in the decision-making process, maximizing system performance, and ensuring the accuracy and reliability of experimental results.
Smart Images

Figure CN120219932A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multimodal fusion, and specifically to an experimental intelligent navigation method and system based on multimodal intention understanding. Background Art
[0002] In today's education system, the chemistry discipline occupies a crucial position. As the core link of chemistry teaching, chemistry experiments play an irreplaceable role in helping students deeply understand chemical theoretical knowledge, cultivate practical abilities and scientific thinking. However, with the continuous development and popularization of chemistry education, the safety issues faced in the operation of chemistry experiments in the teaching process have become increasingly prominent, and have become one of the most challenging problems in the current education field.
[0003] Currently, in order to solve the safety problems in chemistry teaching experiments, intelligent experimental platforms based on virtual-real fusion have been proposed. However, these methods often require that the range and mode of input operations such as the user's gestures and voices must meet certain constraint conditions, that is, it is difficult to meet the application requirements of natural interaction. Secondly, the intelligent experimental process in middle schools is a typical human-machine collaborative interaction process. However, most of the current intelligent experimental application systems are mainly "machine-centered" and it is difficult to achieve the intelligent goal of completing the experimental process through a collaborative mode between humans and machines.
[0004] Therefore, there is an urgent need for an experimental intelligent navigation method and system based on multimodal intention understanding to solve the above problems. Summary of the Invention
[0005] The purpose of the present invention is to provide an experimental intelligent navigation method and system based on multimodal intention understanding, which can achieve the intelligent goal of completing the experimental process through a collaborative mode between humans and machines.
[0006] To achieve the above object, the present invention is realized through the following technical solutions:
[0007] On the one hand, an experimental intelligent navigation method based on multimodal intention understanding is provided, including the following steps:
[0008] S1: Use the fusion of visual modality and sensor modality to obtain the user's intention;
[0009] S2: Execute a multimodal fusion algorithm based on dynamic game according to the user's intention obtained in step S1;
[0010] S3: Construct an experimental assistance model according to the multimodal fusion algorithm in step S2.
[0011] Preferably, in the step S1, the visual modality includes:
[0012] Use an RGB camera to obtain the user's operating object in the experimental scenario, and use a CNN to extract the information features of the RGB to obtain the visual feature information about the visual modality
[0013] The sensor modalities include:
[0014] By setting an attitude sensor and a touch sensor on the user's operating object in the experimental scenario, collect the acceleration and rotation angle of the operating object, determine the user's operation behavior, and use the SrNet network to process the data of the attitude sensor and the touch sensor to obtain the feature information about the sensor modality
[0015] Preferably, the obtained feature information is respectively input into the self-attention module for preprocessing, and is respectively multiplied and transformed with the corresponding learnable projection matrix to obtain According to the obtained and obtain the weight matrix W for this modality α ∈R L×L , and the calculation method is:
[0016]
[0017] The calculation method of the weighted feature vector is:
[0018]
[0019] Among them, is the modality feature obtained after importance amplification.
[0020] Preferably, the processed modality features are vector-concatenated to obtain the complete feature vector F combine ∈R 2L×1 ;
[0021] After the obtained feature vector F combine passes through the fully connected layer and normalization processing, the Softmax activation function is used in the output layer to convert the output into a probability distribution I vs ;
[0022] If the user does not select voice input, then I vs The I with the largest probability value in the probability distribution is the final intention, as follows:
[0023] I = argmax i I vs (i) (3)
[0025] If the user selects voice input, the user intentions obtained from the visual modality and the sensor modality are fused with the intentions obtained from the voice modality.
[0026] Preferably, step S2 is specifically as follows:
[0027] Define the overall utility function U(λ) to measure the quality of the fusion result:
[0028] U(λ) = λ(1 - H(I vs )) + (1 - λ)(1 - H(I a )) (4)
[0029] Where H(I) is the information entropy, λ is the weight of the visual and sensor modalities, 1 - λ is the weight of the voice modality, I(x) is the probability of the experimental intention x occurring under the given experimental conditions. Among them, the calculation process of H(I) is as follows:
[0030] H(I) = -∑ x I(x)logI(x) (5)
[0031] Calculate the gradient of the utility function with respect to the weight λ, and use the gradient to update λ. The formula for calculating the gradient is:
[0032]
[0033] It means that the difference between the information entropy of the visual sensor modality and the information entropy of the voice modality determines the direction of weight adjustment. According to the gradient, the weight update rule is as follows:
[0034] λ ← λ + η(H(I a ) - H(I vs )) (7)
[0036] Where η is the learning rate, a parameter used to control the weight update step size;
[0037] When H(I a ) > H(I vs ) it means that the voice modality is currently unstable or inaccurate, then this rule will increase λ and reduce the weight of the voice modality;
[0038] Based on the dynamically adjusted weights, the final fusion result can be obtained as follows:
[0039] I = max(λI vs + (1 - λ)I a )
[0040] (8).
[0041] Preferably, in step S3, constructing the experimental auxiliary model includes the following steps:
[0042] S31: Define the basic elements of the Bayesian network, including: variable nodes and the states of the nodes;
[0043] S32: Construct the relevant CPT for each node to describe the probability of the state of the node under the specific states of its parent nodes;
[0044] S33: Calculate the posterior probability P(X i |parents(X i )) using the new data and the previous CPT through the Bayesian rule.
[0045] Preferably, in the step S31,
[0046] The variable nodes are the experimental steps of the experimental operation;
[0047] The states of the nodes are defined as "completed" and "not completed".
[0048] Preferably, the step S3 is specifically:
[0049]
[0050] Wherein, X i represents the node, P(parents(X i )|X i )×P(X i ) is the probability value obtained according to the CPT, P(X i ) is the probability value of the obtained user intention, and P(parents(X i )) is calculated through the total probability formula of all relevant combinations.
[0051] On the other hand, a navigation system based on the experimental intelligent navigation method for multi-modal intention understanding as described in claim 1 is provided, including:
[0052] A data acquisition and preprocessing module, configured to: use the fusion of the visual modality and the sensor modality to obtain the user's intention;
[0053] An algorithm execution module, configured to execute a multi-modal fusion algorithm based on dynamic game according to the obtained user intention;
[0054] A model construction module, configured to: construct an experimental assistance model according to the multi-modal fusion algorithm.
[0055] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0056] 1. The present invention realizes the intelligent goal of completing the experimental process through a collaborative mode between humans and machines;
[0057] 2. The present invention uses information entropy to dynamically adjust the weights of various modal data in the decision-making process. Through this method, each modal data adjusts its influence in the decision-making, which can maximize the performance of the entire system.
[0058] 3. Through the fusion of experimental data, the present invention can be adjusted in a changing environment and user intentions, ensuring the accuracy and reliability of the experimental results.
[0059] 4. The model in the present invention will process the sequence errors of user operations according to the specified experimental rules, provide real-time feedback and guidance, and will also evaluate whether the user's operation behavior is safe, reminding the user to observe the experimental phenomena, etc. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 is the flowchart of the method of the present invention;
[0061] Figure 2 is the Bayesian network diagram of the "sodium-water reaction" experiment in the embodiment of the present invention;
[0062] Figure 3 is the schematic diagram of the system structure of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0063] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by this application.
[0064] In the present invention, terms such as "upper", "lower", "left", "right", "front", "rear", "vertical", "horizontal", "side", "bottom", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. They are only relationship words determined to facilitate the description of the structural relationship of each component or element of the present invention, and do not specifically refer to any component or element of the present invention, and should not be construed as a limitation to the present invention.
[0065] In the present invention, terms such as "fixed connection", "connected", "connected" should be understood in a broad sense, which can mean a fixed connection, an integral connection or a detachable connection; it can be directly connected or indirectly connected through an intermediate medium. For those skilled in the relevant scientific research or technology in this field, the specific meanings of the above terms in the present invention can be determined according to specific circumstances, and should not be construed as a limitation to the present invention.
[0066] Embodiment:
[0067] Such as Figure 1As shown in the figure, this embodiment provides an experimental intelligent navigation method based on multimodal intention understanding, including the following steps: S1: Use the fusion of visual modality and sensor modality to obtain the user's intention;
[0068] S2: According to the user intention obtained in step S1, execute a multimodal fusion algorithm based on dynamic game;
[0069] S3: According to the multimodal fusion algorithm in step S2, construct an experimental assistance model.
[0070] In step S1, the visual modality includes:
[0071] Use a head-mounted RGB camera to obtain the user's operation object in the experimental scene, and use CNN to extract the information features of RGB to obtain the visual feature information about the visual modality to understand the experimental scene;
[0072] The sensor modality includes:
[0073] By setting an attitude sensor and a touch sensor on the user's operation object (in this embodiment, the operation object is specifically a beaker) in the experimental scene, used to sense the acceleration and rotation angle of the beaker, determine the user's operation behaviors such as "pouring", "clamping", "dropping", etc., and use the SrNet network to process the data of the attitude sensor and the touch sensor to obtain the feature information about the sensor modality
[0074] Then the obtained feature information is respectively input into the self-attention module for preprocessing, and (α = v, s; L represents the dimension of the feature) is respectively multiplied by the corresponding learnable projection matrix to obtain Then according to the obtained and According to formula (1), the weight matrix W of this modality can be obtained α ∈R L×L :
[0075]
[0076] Then, the final weighted feature vector is obtained by formula (2):
[0077]
[0078] Among them, is the modality feature obtained after importance amplification, that is, the self-attention mechanism can obtain more important feature information, and the computational complexity will be smaller;
[0079] After processing the feature information of each modality, in order to obtain the complete user intention, in this embodiment, the processed modality features are concatenated into vectors to obtain a complete feature vector F combine ∈R 2L×1 ; Then the obtained feature vector F combine After passing through the fully connected layer and normalization processing, the Softmax activation function is used in the output layer to convert the output into a probability distribution I vs ;
[0080] If the user does not select voice input, then the I vs The I with the largest probability value in the probability distribution is the final intention, as shown in formula (3):
[0081] I = argmax i I vs (i) (3)
[0083] If the user selects to input voice information, the system needs to fuse the intentions obtained from the visual and sensor modalities with the intention obtained from the voice modality to further confirm the user's true intention. Therefore, this embodiment adopts the ideological method of dynamic game and dynamically adjusts the weights between modalities based on information entropy to obtain the user's final intention.
[0084] In step S2, information entropy is used as the basis for adjusting the modality weights, and the overall utility function U(λ) is defined to measure the quality of the fusion result, as shown in formula (4):
[0085] U(λ) = λ(1 - H(I vs )) + (1 - λ)(1 - H(I a )) (4)
[0087] Among them, H(I) is the information entropy, which can be calculated by formula (5):
[0088] H(I) = -∑ x I(x)logI(x) (5)
[0090] λ is the weight of the visual and sensor modalities, 1 - λ is the weight of the voice modality, I(x) is the probability of the occurrence of the experimental intention x under the given experimental conditions. By maximizing U(λ), a weight configuration is found to maximize the fusion of modality information and improve the accuracy;
[0091] To dynamically adjust the weights, it is necessary to calculate the gradient of the utility function with respect to the weight λ and use it to update λ. The calculation of the gradient is as shown in formula (6):
[0092]
[0093] It is shown that the difference between the information entropy of the visual sensor modality and the information entropy of the speech modality determines the direction of weight adjustment. According to the gradient, the weight update rule is as shown in Equation (7):
[0094] λ←λ+η(H(I a )-H(I vs )) (7)
[0096] where η is the learning rate, a parameter that controls the weight update step size. When H(I a ) is greater than H(I vs ), it indicates that the speech modality may be relatively unstable or inaccurate at present. Then this rule will increase λ and reduce the weight of the speech modality, that is, the modality with lower information entropy will obtain a greater weight, and vice versa;
[0097] Based on the dynamically adjusted weights, the final fusion result can be obtained, as shown in Equation (8):
[0098] I=max(λI vs +(1-λ)I a ) (8)
[0100] Through this method, the fusion of experimental data can be adjusted in a changing environment and user intentions to ensure the accuracy and reliability of experimental results.
[0101] In the traditional experimental process, a teacher often needs to give different guidance and hints for the operation behaviors of students. However, in the case of a teacher tutoring multiple students, it is inevitable that the teacher will be overstretched. Therefore, in response to this situation, this embodiment makes an innovation to the previous intelligent experiment by adding an intelligent experiment assistance model. Taking the sodium-water reaction experiment as an example, this embodiment implements the intelligent experiment assistance model based on the Bayesian network according to the experimental rules.
[0102] The construction of the experiment assistance model includes the following steps:
[0103] (1) Define the basic elements of the Bayesian network, including variable nodes and the states of the nodes. According to prior knowledge, the variable nodes can be defined as the necessary experimental steps of the experiment operation, and can also include "experimental safety", "observation results", etc. The states of the nodes are divided into "completed" and "not completed". According to the key steps of the sodium-water reaction experiment, a Bayesian network is created, as Figure 2 shown. In the initial state, the states of all nodes are not completed;
[0104] (2) According to the designed Bayesian network, construct the relevant CPT for each node to describe the probability of the state of the node under the specific state of its parent node;
[0105] For example, if node B depends on node A, the construction of the conditional probability table CPT(B|A) for node B is shown in the CPT table of node B in Table 1:
[0106] Table 1 CPT table of node B
[0107] A P(B = True|A) P(B = False|A) True t 1-t False f 1-f
[0108] Where t and f are weights set by experts based on experience, representing the probability of B being completed when A is completed or not. If node B has multiple parent nodes, then when constructing the CPT table, it is necessary to consider the influence of the combination of all parent node states on the state of B;
[0109] (3) When the user completes an experimental step, the input result will update the state of the node of this step. And for each experimental step, the model can check whether all its prerequisites (i.e., parent nodes) are marked as completed. If not, the model will find that the experimental step order is incorrect, warn the user and provide specific operation guidance according to the dependency relationship of the Bayesian network;
[0110] In addition, in the Bayesian network, the model will use new data and the previous CPT to calculate the posterior probability through the Bayesian rule. This process can help the model dynamically predict and adjust the prediction of the danger of the experimental operation and remind the user to observe the experimental phenomenon. The posterior probability P(X i |parents(X i )) of the state of i can be obtained from formula (9):
[0111]
[0112] Where P(parents(X i )|X i )×P(X i ) is the probability value obtained according to the CPT, P(X i ) is the probability value of the user's intention obtained according to, and P(parents(X i )) is calculated through the total probability formula of all relevant combinations. Through this method, the model will continuously update the network state according to the user's experimental operation results, so as to improve the accuracy of guidance and prediction.
[0113] With this method based on Bayesian networks, the model can determine in real time whether the experimental operation sequence is correct according to the parent nodes of the current node, and specifically guide the user based on the dependency relationship. At the same time, for steps involving potential risks, the model evaluates the safety risks according to the posterior probabilities of relevant nodes. If the probability value is higher than the threshold, the model will issue relevant warnings. For key observation points, the model will remind the user to pay attention to observing and recording specific experimental phenomena, such as color changes and bubble generation, at appropriate times to ensure the accurate recording and analysis of experimental data.
[0114] As Figure 3 shown, this embodiment also provides a navigation system based on the above experimental intelligent navigation method based on multi-modal intention understanding, including:
[0115] A data acquisition and preprocessing module for: using the fusion of visual modality and sensor modality to obtain the user's intention;
[0116] An algorithm execution module for executing a multi-modal fusion algorithm based on dynamic game according to the obtained user intention;
[0117] A model construction module for: constructing an experimental assistance model according to the multi-modal fusion algorithm.
[0118] The above is a specific description of the preferred embodiment of the present invention, but the present invention is not limited to the described embodiment. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.
Claims
1. An experimental intelligent navigation method based on multimodal intent understanding, characterized in that: The following steps are involved: S1: Using visual modality and sensor modality fusion to obtain user intention; S2: Execute a multimodal fusion algorithm based on dynamic game according to the user intention obtained in step S1; S3: Construct an experimental auxiliary model according to the multimodal fusion algorithm in step S2.
2. According to claim 1, an experimental intelligent navigation method based on multimodal intent understanding is characterized in that: In step S1, the visual modality includes: Use an RGB camera to obtain the user's operating object in the experimental scene, and use CNN to extract RGB information features to obtain visual feature information about the visual modality. Sensor modalities include: By setting gesture sensors and touch sensors on the user's operating object in the experimental scene, the acceleration and rotation angle of the operating object are collected to determine the user's operating behavior, and the SrNet network is used to process the gesture sensors and touch sensors to obtain characteristic information about the sensor mode.
3. According to claim 2, an experimental intelligent navigation method based on multimodal intention understanding is characterized in that: The feature information to be obtained They are input into the self-attention module for preprocessing. Multiply and transform with the corresponding learnable projection matrix to obtain According to the obtained and Get the weight matrix W about this mode α ∈R L×L , calculated as: The weighted eigenvector is calculated as: in, It is the modal feature obtained after importance amplification.
4. According to claim 3, an experimental intelligent navigation method based on multimodal intention understanding is characterized in that: The processed modal features are vector-joined to obtain the complete feature vector F combine ∈R 2L×1 ; The obtained feature vector F combine After the full connection layer and normalization, the Softmax activation function is used in the output layer to convert the output into a probability distribution I vs ; If the user does not select voice input, I vs The I with the largest probability value in the probability distribution is the final intention, as shown below: If the user chooses voice input, the user intent obtained by the visual modality and sensor modality is merged with the intent obtained by the voice modality.
5. According to claim 1, an experimental intelligent navigation method based on multimodal intent understanding is characterized in that: The step S2 is specifically: Define the overall utility function U(λ) to measure the quality of the fusion result: U(λ)=λ(1-H(I vs ))+(1-λ)(1-H(I a ))(4) Where H(I) is the information entropy, λ is the weight of the visual and sensor modalities, 1-λ is the weight of the speech modality, and I(x) is the probability of the experimental intention x occurring under given experimental conditions. The calculation process of H(I) is as follows: H(I)=-∑ x I(x)logI(x)(5) Calculate the gradient of the utility function with respect to the weight λ and use the gradient to update λ. The gradient is calculated as follows: The difference between the information entropy of the visual sensor modality and the information entropy of the speech modality determines the direction of weight adjustment. According to the gradient, the weight update rule is as follows: Among them, η is the learning rate, which is a parameter used to control the step size of weight update; When H(I a )>H(I vs ), it means that the speech mode is currently unstable or inaccurate, then the rule will increase λ and reduce the weight of the speech mode; The final fusion result can be obtained based on the dynamically adjusted weights, as shown below: I=max(λI vs +(1-λ)I a ) (8).
6. The experimental intelligent navigation method based on multimodal intention understanding according to claim 1, characterized in that: In step S3, constructing the experimental auxiliary model includes the following steps: S31: Define the basic elements of Bayesian networks, including variable nodes and node states; S32: construct a related CPT for each node to describe the probability of the node state under the specific state of its parent node; S33: Use the new data and the previous CPT to calculate the posterior probability P(X i |parents(X i )).
7. The experimental intelligent navigation method based on multimodal intention understanding according to claim 6 is characterized in that: In the step S31, The variable nodes are the experimental steps of the experimental operation; The status of a node is defined as "completed" or "uncompleted".
8. The experimental intelligent navigation method based on multimodal intention understanding according to claim 6 is characterized in that: The step S3 is specifically: Among them, X i Represents a node, P(parents(X i )|X i )×P(X i ) is the probability value obtained according to CPT, P(X i ) is the probability value based on the obtained user intention, P(parents(X i )) is calculated using the total probability formula for all relevant combinations.
9. A navigation system based on the experimental intelligent navigation method based on multimodal intention understanding as claimed in claim 1, characterized in that: include: Data acquisition and preprocessing module, used to: use visual modality and sensor modality fusion to obtain user intention; An algorithm execution module is used to execute a multimodal fusion algorithm based on dynamic game according to the acquired user intention; The model building module is used to build an experimental auxiliary model based on the multimodal fusion algorithm.
Citation Information
Cited By
Image processing method and system for multi-view experimental operation analysis
CN121121413A