Self-adaptive multi-mode complementary intention understanding method, system and equipment

Through the adaptive multimodal complementary intention understanding method, using gesture, speech and image modal information, combined with knowledge graphs and mixed expert models, the problem of difficult understanding of modal information in the elderly is solved in the existing technology, and more accurate and intuitive human-computer interaction is achieved.

CN120217294APending Publication Date: 2025-06-27SHANDONG XIEHE UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510290457.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Existing human-computer interaction technology is difficult to accurately understand the modal information of the elderly, resulting in obstacles in the interaction process and the real needs of the elderly are not accurately understood.

Method used

Adaptive multimodal complementary intention understanding method is adopted to build a knowledge graph by obtaining gesture, speech and image modal information in real time, and multimodal intention fusion is carried out using a complementary attention mechanism based on knowledge graph and a hybrid expert model.

Benefits of technology

It improves the accuracy of intention understanding, creates a more intuitive and responsive interactive process, enriches the interactive experience, and is especially suitable for the elderly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217294A_ABST
    Figure CN120217294A_ABST
Patent Text Reader

Abstract

The invention discloses a self-adaptive multi-mode complementary intention understanding method, system and device, and mainly relates to the technical field of multi-mode complementary intention understanding. Comprising the following steps: acquiring a gesture feature vector sequence in a man-machine interaction process in real time, and processing the gesture feature vector sequence; acquiring a continuous audio stream in real time, and segmenting the continuous audio stream into pause or fixed time windows based on voice; building an experimental environment, and obtaining a real-time image from the experimental environment as one of the modalities; constructing a knowledge graph according to the acquired gesture modality, voice modality and image modality; carrying out multi-modal intention fusion extraction by adopting a complementary attention mechanism based on a knowledge graph for the voice mode and the image mode; and carrying out multi-modal intention fusion by using a hybrid expert model. The method has the beneficial effects that the intention understanding accuracy is improved, and the interaction experience is enriched by creating a more intuitive and more rapid-response interaction process between the robot and the elderly users thereof.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multimodal complementary intention understanding, and particularly to an adaptive multimodal complementary intention understanding method, system and device. Background Art

[0002] With the rapid development of artificial intelligence technology, the interaction between humans and robots has become increasingly frequent and important. Natural and intelligent interaction with robots has become one of the key directions for the development of this field. However, in practical applications, especially when involving the elderly population, there are many problems that need to be solved urgently.

[0003] At present, although the existing human-computer interaction technologies can achieve information transmission and task execution to a certain extent, they expose obvious deficiencies when facing the special group of the elderly. Due to the natural degradation of physical functions, the expression of their modal information (such as speech, gestures, expressions, etc.) is often not clear and accurate enough. For example, the elderly may have slurred speech due to the decline of the pronunciation organ function; due to the decrease in limb flexibility, the gesture expressions are not standardized or difficult to be accurately recognized; the facial expressions may also become blurred due to the weakening of muscle control ability. These factors have brought great difficulties to the robot to obtain accurate executable intentions.

[0004] Most of the existing human-computer interaction systems are designed based on relatively idealized interaction models, with high requirements for the accuracy and integrity of modal information. When facing the unclear modal information input from the elderly, these systems are often unable to effectively understand and process it, resulting in obstacles in the interaction process, unable to accurately understand the real needs of the elderly, and thus difficult to provide appropriate services and assistance. This makes it difficult to truly achieve natural and intelligent interaction between humans and robots among the elderly population, severely restricting the application and development of artificial intelligence technology in this field and becoming one of the main bottlenecks in the current development of artificial intelligence.

[0005] Therefore, there is an urgent need for an adaptive multimodal complementary intention understanding method, system and device to solve the above problems. Summary of the Invention

[0006] The purpose of the present invention is to provide an adaptive multimodal complementary intention understanding method, system and device, which not only improves the accuracy of intention understanding, but also enriches the interaction experience by creating a more intuitive and responsive interaction process between the robot and its elderly users.

[0007] To achieve the above object, the present invention is realized through the following technical solutions:

[0008] On the one hand, an adaptive multimodal complementary intention understanding method is provided, including the following steps:

[0009] S1: Obtain the gesture feature vector sequence in real time during human-computer interaction, and process the gesture feature vector sequence;

[0010] S3: Set up an experimental environment, and obtain real-time images from the experimental environment as one of the modalities;

[0011] S3: Set up an experimental environment, and obtain real-time images from the experimental environment as one of the modalities;

[0012] S4: Construct a knowledge graph based on the gesture modality, speech modality, and image modality obtained in steps S1 - S3;

[0013] S5: Adopt a complementary attention mechanism based on the knowledge graph for multi-modal intention fusion extraction for the speech modality and image modality;

[0014] S6: Use a mixture of experts model for multi-modal intention fusion.

[0015] Preferably, the step S1 includes the following steps:

[0016] S11: Extract the feature sequence H = {h1, h2,..., h N} of the gesture features, and use the learned weight matrices W K and W V . Only calculate the key K and value V matrices once in the first output vector. The calculation formula is as follows:

[0017] K = HW K , V = HW V (1);

[0018] S12: For each iteration or self-attention layer, recalculate the query matrix Q using the output of the previous layer and the learned weight matrix W Q . The calculation formula is as follows:

[0019] Q = HW Q (2)

[0020] where H is updated in each iteration to reflect the newly obtained features;

[0021] S13: Calculate the attention scores and output as follows:

[0022]

[0023] H′ = AV(4)

[0024] where A is the attention score matrix, H′ is the updated feature vector sequence calculated using the self-attention mechanism, and d k is the dimension of the key vector;

[0025] S14: After the self-attention mechanism calculation, apply layer normalization and residual connection to obtain the final set of feature vectors, and the calculation formula is as follows:

[0026] H" = LN(H + H′)(5).

[0027] Preferably, the steps S2 - S3 are specifically: for the speech and image modalities, obtain the speech feature sequence F and the image feature sequence X during the feature extraction stage, and perform the self-attention mechanism calculation process as in steps S11 - S14 to obtain the speech feature vector set F" and the image feature vector set X".

[0028] Preferably, the step S4 includes:

[0029] S41: Identify entities in the environment and analyze the spatial relationships between the detected entities;

[0030] S42: Create nodes for each identified entity and edges for each relationship. Each unique entity has a node, and each edge can be labeled with the relationship type it represents;

[0031] S43: Use a data structure that supports efficient querying and updating to implement the knowledge graph.

[0032] Preferably, the step S5 includes:

[0033] S51: For the given speech feature vector F_speech and the image feature vector X_image, calculate Q, K, and V:

[0034]

[0035]

[0036] Where, and are trainable weight matrices specific to each type of feature vector, and are used to calculate Q, K, and V respectively;

[0037] S52: Perform self-attention calculation to obtain a weighted feature vector that highlights the complementary aspects of speech and image data:

[0038]

[0039] F A = AV (10)

[0041] Where, d k is the dimension of the key vector, which is used to ensure that the dot product is properly scaled;

[0042] S53: Apply residual connection and layer normalization to integrate the original and attention-weighted feature vectors:

[0043] F′ A = LN(F A + A) (11)

[0045] where F′ A represents the modal feature vector updated after one iteration of the complementary attention mechanism;

[0046] S54: Combine the fusion loss and the consistency loss, and add a regularization term:

[0047]

[0048] where λ f , λ kg , λ reg are hyperparameters that balance each component in the loss function.

[0049] Preferably, the step S6 includes:

[0050] S61: The gating network in the mixture-of-experts model dynamically determines the contribution of each expert according to the input features, and the calculation formula is:

[0051] S′ = S combined [H", F", X", F′ A (13)

[0053] G = softmax(W g S′ + b g ) (14)

[0055] where G is the weight vector of each expert, W g is the weight matrix, S′ is the concatenation of the feature vectors of all modalities processed by their respective self-attention mechanisms and expert modules, and b g is the bias term;

[0056] S62: The mixture-of-experts model calculates the weighted sum of the expert outputs and performs an additional normalization step:

[0057]

[0058] where E i (F i ′) represents the output of the i th th expert module after applying self-attention to its input feature F i ′, and G i ​It is the weight assigned by the gating network to this expert. After obtaining the final fused feature, it is input into a multi-layer perceptron, and connected to the softmax layer to obtain the final intention classification result.

[0059] On the other hand, an adaptive multi-modal complementary intention understanding system is provided, including:

[0060] A data acquisition module, used for: obtaining the gesture feature vector sequence in real time during the human-computer interaction process, and processing the gesture feature vector sequence; obtaining the continuous audio stream in real time and segmenting it into based on pauses in speech or fixed time windows; setting up an experimental environment, and obtaining real-time images from the experimental environment as one of the modalities;

[0061] A knowledge graph construction module, used for: constructing a knowledge graph according to the obtained gesture modality, speech modality, and image modality;

[0062] A multi-modal fusion module, used for: adopting a complementary attention mechanism based on the knowledge graph for multi-modal intention fusion extraction for the speech modality and the image modality; using a mixture of experts model for multi-modal intention fusion.

[0063] On the other hand, an adaptive multi-modal complementary intention understanding device is provided, including a processor and a memory for storing a computer program. When the processor executes the computer program, the steps of the method according to any one of claims 1-6 are implemented.

[0064] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0065] 1. An innovative cross-modal mixture of experts intention fusion model is introduced. This model is equipped with a self-attention mechanism, which can synthesize and interpret ambiguous or incomplete inputs in multiple modalities, thereby improving the robot's ability to understand and effectively respond to complex user requirements;

[0066] 2. The image-audio modality complementary attention model based on the knowledge graph is pioneered. This model utilizes the inherent synergy between visual and auditory data. This method not only improves the accuracy of intention understanding, but also enriches the interaction experience by creating a more intuitive and responsive interaction process between the robot and its elderly users. Description of the Drawings

[0067] Figure 1 is the flowchart of the method of the present invention;

[0068] Figure 2 is the flowchart of the speech-image modality complementary attention mechanism of the present invention;

[0069] Figure 3 is the flowchart of the mixture of experts multi-modal intention fusion model of the present invention;

[0070] Figure 4 It is a schematic diagram of the system structure of the present invention. Specific embodiments

[0071] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by this application.

[0072] In the present invention, terms such as "upper", "lower", "left", "right", "front", "rear", "vertical", "horizontal", "side", "bottom", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. They are only relational terms determined for the convenience of describing the structural relationship of each component or element of the present invention, and do not specifically refer to any component or element of the present invention, and should not be construed as a limitation to the present invention.

[0073] In the present invention, terms such as "fixed connection", "connected", "connected" should be understood in a broad sense, which may mean a fixed connection, an integral connection or a detachable connection; it may be directly connected or indirectly connected through an intermediate medium. For those skilled in relevant scientific research or technology in this field, the specific meanings of the above terms in the present invention can be determined according to specific circumstances, and should not be construed as a limitation to the present invention.

[0074] Embodiment:

[0075] As Figure 1 shown, this embodiment provides an adaptive multi-modal complementary intention understanding method, including the following steps:

[0076] S1: Real-time obtain the gesture feature vector sequence during the human-computer interaction process, and process the gesture feature vector sequence;

[0077] S2: Real-time obtain a continuous audio stream and segment it into based on pauses in the speech or fixed time windows;

[0078] S3: Build an experimental environment and obtain real-time images from the experimental environment as one of the modalities;

[0079] S4: Construct a knowledge graph according to the gesture modality, speech modality and image modality obtained in steps S1-S3;

[0080] S5: Adopt a complementary attention mechanism based on the knowledge graph for multi-modal intention fusion extraction for the speech modality and the image modality;

[0081] S6: Use a mixture of experts model for multi-modal intention fusion.

[0082] In view of the behavioral characteristics of the elderly, in this embodiment, a self-attention mechanism is adopted in the acquisition of modal intents. Through the calculation of modal feature attention scores, the fuzzy or incomplete inputs of the elderly are captured, and the effective features in the modal features are highlighted, thereby improving the accuracy of the robot in obtaining the intents of the elderly.

[0083] For the feature sequence H = {h i , h2, …, h N} of gesture feature extraction, the learned weight matrices W K and W V are used. The key (K) and value (V) matrices are only calculated once in the first output vector, and the calculation formula is as shown in (1).

[0084] K = HW K , V = HW V (1)

[0085] For each iteration or self-attention layer, the output of the previous layer (initially H) and the learned weight matrix W Q are used to recalculate the query matrix (Q), and the calculation formula is as follows.

[0086] Q = HW Q (2)

[0087] Among them, H is updated in each iteration to reflect the newly obtained features.

[0088] The attention scores and outputs are calculated as follows:

[0089]

[0090] H′ = AV(4)

[0091] Among them, A is the attention score matrix, H′ is the updated feature vector sequence calculated by applying the self-attention mechanism, and d k is the dimension of the key vector.

[0092] After the self-attention mechanism calculation, layer normalization and residual connection are applied to obtain the final set of feature vectors, and the calculation formula is as shown in (5):

[0093] H" = LN(H + H′)(5)

[0094] In this embodiment, for the speech and image modalities, the speech feature sequence F and the image feature sequence X are obtained in the feature extraction stage, and the same self-attention mechanism calculation as that of the gesture feature is performed on both, obtaining the speech feature vector set F" and the image feature vector set X". Since the calculation process is similar (only the parameters are different), the self-attention calculation processes of speech and image are not described one by one.

[0095] In this embodiment, a relevant knowledge graph is created for the specific application scenario of the elderly and robots collaborating to build blocks. The creation of the knowledge graph requires determining the key entities and their potential relationships in the task context. The entities specifically include building blocks (color, shape), actions (pick up, place, stack), participants (the elderly, robots), and locations (table location, block location). The relationships between entities include spatial relationships (above, beside, below), temporal relationships (before, after), action relationships (pick up, place, stack), and descriptive relationships (color, shape);

[0096] The construction of the knowledge graph involves populating it with instances of these entities and their relationships based on observed interactions. Data collection collects data from voice inputs (commands, descriptions of the elderly) and real-time environmental images (the positions and states of the building blocks, actions taken by humans and robots). Entities (building blocks) and actions are extracted from the voice inputs, and the language structure is parsed to determine relationships (e.g., "place building block A on top of building block B"). First, the entities in the environment are identified, and the spatial relationships between the detected entities are analyzed. Nodes are created for each identified entity, and edges are created for each relationship. Each unique entity (building block, action, participant) has a node, and each edge can be labeled with the type of relationship it represents. A data structure that supports efficient querying and updating is used to implement the knowledge graph because real-time interactions may require dynamic changes to the graph. In this paper, a graph database (Neo4j) is used, which provides built-in mechanisms for relationship querying and pattern matching.

[0097] Considering the complementarity between multiple modalities, in order to obtain the complementarity between modalities and further improve intent understanding, in this embodiment, a complementary attention mechanism based on the knowledge graph is adopted for multi-modal intent fusion extraction for the voice modality and the image modality. The obtained modality features are converted into a compatible vector format and fed into the complementary attention mechanism. As Figure 2 shown, in the complementary attention mechanism, the complementary relationship between the two modalities is obtained through the knowledge graph and used as a condition for attention calculation. The specific process is as follows:

[0098] For the given voice feature vector F_voice and image feature vector X_image, calculate Q, K, and V.

[0099]

[0100] Among them, and are trainable weight matrices specific to each type of feature vector, and are used to calculate Q, K, and V respectively.

[0101] Perform self-attention calculation to obtain a weighted feature vector that highlights the complementary aspects of the voice and image data.

[0102]

[0103] F A = AV(10)

[0104] where d k is the dimension of the key vector, ensuring that the dot product is appropriately scaled.

[0105] Residual connections and layer normalization are applied to integrate the original and attention-weighted feature vectors, improving stability and performance.

[0106] F′ A = LN(F A + A)(11)

[0107] F′ A represents the modality feature vector updated after one iteration of the complementary attention mechanism.

[0108] To ensure that the fused features are consistent with the semantic relationships encoded in the knowledge graph, in addition to the fusion loss, this paper adopts the knowledge graph consistency loss Combining the fusion loss and the consistency loss, a regularization term is added to prevent overfitting, and the specific formula is shown in (12).

[0109]

[0110] where λ f , λ kg , λ reg are hyperparameters that balance each component in the loss function.

[0111] In the intention fusion stage, this embodiment designs a Mixture of Experts (MoE) model to achieve this goal. As Figure 3 shown, separate expert modules are designed for each modality (speech, gesture, image) within the MoE framework. Each expert module focuses on capturing patterns, relationships, or characteristics specific to its corresponding modality:

[0112] The gating network in the Mixture of Experts model dynamically determines the contribution of each expert based on the input features. It outputs a set of weights indicating how much each expert should contribute to the final prediction. The calculation formulas are shown in (13), (14):

[0113] S′ = S combined [H", F", X", F′ A #(13)

[0114] G = softmax)W g S′ + b h )#(14)

[0115] Among them, G is the weight vector of each expert, and W g is the weight matrix, S′ is the concatenation of the feature vectors processed by all modalities through their respective self-attention mechanisms and expert modules, and b g is the bias term.

[0116] After determining the weights of each expert, the MoE model calculates the weighted sum of the expert outputs and then performs an additional normalization step to produce the final output:

[0117]

[0118] Among them, E i (F i ′) represents the output of the i th -th expert module after applying self-attention to its input feature F i ′, and G i is the weight assigned to this expert by the gating network. After obtaining the final fused feature, it is input into a multi-layer perceptron (MLP), and finally connected to a softmax layer to obtain the final intent classification result.

[0119] In the mixture-of-experts model, the following loss function is designed.

[0120]

[0121] Among them, is the loss of a specific task, is the regularization term of the i-th expert, and λ is the hyperparameter for balancing regularization.

[0122] The above is a specific description of the preferred embodiment of the present invention, but the present invention is not limited to the described embodiment. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.

Claims

1. An adaptive multimodal complementary intention understanding method, characterized in that: The following steps are involved: S1: Acquire gesture feature vector sequences in real time during human-computer interaction and process the gesture feature vector sequences; S2: Acquire a continuous audio stream in real time and segment it into pauses or fixed time windows in speech; S3: Build an experimental environment and obtain real-time images from the experimental environment as one of the modalities; S4: construct a knowledge graph according to the gesture modality, voice modality, and image modality obtained in steps S1-S3; S5: A complementary attention mechanism based on knowledge graph is used for multimodal intent fusion extraction for speech and image modalities; S6: Use hybrid expert model for multimodal intent fusion.

2. According to claim 1, the adaptive multimodal complementary intention understanding method is characterized in that: The step S1 comprises the following steps: S11: Extract gesture features and perform feature sequence H = {h1, h2, ..., h N }, using the learned weight matrix W K and W V , the key K and value V matrices are calculated only once in the first output vector, and the calculation formula is as follows: K=HW K ,V=HW V (1); S12: For each iteration or self-attention layer, use the output of the previous layer and the learned weight matrix W Q Recalculate the query matrix Q, the calculation formula is as follows: Q=HW Q (2) Among them, H is updated in each iteration to reflect the newly acquired features; S13: Calculate the attention score and output as follows: H′=AV (4) Where A is the attention score matrix, H′ is the updated feature vector sequence after applying the self-attention mechanism, and d k is the dimension of the key vector; S14: After the self-attention mechanism is calculated, layer normalization and residual connection are applied to obtain the final feature vector set, which is calculated as follows: H"=LN(H+H') (5).

3. According to claim 2, the adaptive multimodal complementary intention understanding method is characterized in that: The steps S2-S3 are specifically as follows: for the speech and image modalities, the speech feature sequence F and the image feature sequence X are obtained in the feature extraction stage, and the self-attention mechanism calculation process such as steps S11-S14 is performed to obtain the speech feature vector set F" and the image feature vector set X".

4. According to claim 1, the adaptive multimodal complementary intention understanding method is characterized in that: The step S4 comprises: S41: Identify entities in the environment and analyze and detect the spatial relationships between the entities; S42: Create a node for each identified entity and an edge for each relationship. Each unique entity has a node and each edge can be labeled with the type of relationship it represents. S43: Implement knowledge graphs using data structures that support efficient query and update.

5. The adaptive multimodal complementary intention understanding method according to claim 1, characterized in that: The step S5 comprises: S51: Given a speech feature vector Fspeech and an image feature vector Ximage, calculate Q, K and V: in, and is a trainable weight matrix specific to each type of feature vector, used to calculate Q, K, and V respectively; S52: Perform self-attention calculation to obtain a weighted feature vector that highlights the complementary aspects of speech and image data: F A =OFF (10) Among them, d k is the dimension of the key vector, used to ensure that the dot product is properly scaled; S53: Apply residual connections and layer normalization to integrate the original and attention-weighted feature vectors: F′ A =LN(F A +A) (11) Among them, F′ A Represents the updated modal feature vector after one iteration of the complementary attention mechanism; S54: Combine fusion loss and consistency loss and add regularization terms: Among them, λ f ,λ kg ,λ reg are the hyperparameters for each component in the balanced loss function.

6. The adaptive multimodal complementary intention understanding method according to claim 1, characterized in that: The step S6 comprises: S61: The gating network in the hybrid expert model dynamically determines the contribution of each expert based on the input features, and the calculation formula is: S′=S combined [H",F",X",F′ A ] (13) G=softmax(W g S′+b g ) (14) Among them, G is the weight vector of each expert, W g is the weight matrix, S′ is the concatenation of the feature vectors of all modalities after being processed by their own self-attention mechanisms and expert modules, and b g is the bias term; S62: The hybrid expert model calculates the weighted sum of the expert outputs and performs an additional normalization step: Among them, E i (F′ i ) represents the i-th th The expert module inputs the feature F′ i The output after applying self-attention, G i It is the weight assigned to the expert by the gating network. After obtaining the final fusion feature, it is input into the multi-layer perceptron and connected to the softmax layer to obtain the final intent classification result.

7. An adaptive multimodal complementary intention understanding system, characterized in that: include: The data acquisition module is used to: obtain the gesture feature vector sequence in the human-computer interaction process in real time and process the gesture feature vector sequence; obtain the continuous audio stream in real time and divide it into pauses or fixed time windows based on the speech; build an experimental environment and obtain real-time images from the experimental environment as one of the modalities; The knowledge graph construction module is used to: construct a knowledge graph based on the acquired gesture modality, voice modality, and image modality; The multimodal fusion module is used to: use a complementary attention mechanism based on knowledge graph to perform multimodal intent fusion extraction for speech modality and image modality; and use a hybrid expert model to perform multimodal intent fusion.

8. An adaptive multimodal complementary intention understanding device, characterized in that: The method comprises a processor and a memory for storing a computer program, wherein when the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.