Interactive intention understanding and fast learning system based on multi-mode information fusion

By adopting a key information extraction architecture based on large language models and multi-mode human-computer interaction technology in the robot system, the accuracy of intention understanding in complex daily life scenarios is solved, efficient and accurate task understanding and execution are achieved, and robot recognition capabilities and task success rate are improved.

CN119940369APending Publication Date: 2025-05-06BEIJING UNIV OF TECH

Patent Information

Application Number
CN202510095336.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art is difficult to accurately understand people's interaction intentions in complex daily life scenarios, especially when language interactions may be errors or visual recognition errors, resulting in failure in task execution and reducing the success rate of understanding of interactive intentions.

Method used

The key information extraction architecture based on the large language model is adopted, combined with voice, touch screen, and visual multi-mode interaction technology, semantic key information is extracted through the BERT model and matched with the visual object detection results. The touch screen interaction interface is designed using PyQt for visual display and interactive confirmation, correct errors, and send correction information to the rapid learning system.

Benefits of technology

It improves the robot's intention understanding success rate in complex daily life scenarios, ensures the correct understanding of long-sequence multi-task instructions and high time efficiency, and enhances the robot's recognition ability and task execution success rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940369A_ABST
    Figure CN119940369A_ABST
Patent Text Reader

Abstract

The invention discloses an interactive intention understanding and rapid learning system based on multi-mode information fusion, which aims at semantic understanding of an operation instruction of a robot for picking and placing an article, constructs a key information extraction framework based on a large language model in a task semantic understanding module, and comprises multiple key position and action text information extraction, and multi-mode interactive fusion is carried out on the extracted information, vision and a touch screen, so that correct understanding of the instruction intention is realized. The key information extraction architecture based on the big language model is based on a BERT big language model, and a language processing model is trained through the BERT big language model; the language processing model comprises a long-sequence multi-task instruction and a human instruction data set of mood, scene, positive-sequence and inverse-sequence request instructions, and realizes input of long-sequence multi-task instruction statements and output of a key action sequence; according to the system, efficient and accurate human intention understanding can be realized through a voice, touch screen and vision multi-mode human-computer interaction technology, so that the robot can quickly learn new articles.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of human-computer interaction technology, and provides an intention understanding and rapid learning system based on multi-mode human-computer interaction fusion. Aiming at the field of robot operation, the system can realize efficient and accurate human intention understanding through voice, touch screen, and visual multi-mode human-computer interaction technology, and can integrate language and visual information through touch screen interaction, so that the robot can quickly learn new objects. Background Art

[0002] Electric wheelchairs equipped with robotic arms can provide daily life assistance and improve their ability to take care of themselves, and have broad application prospects. The interactive ability of arm-carrying wheelchairs is crucial, requiring the human-computer interaction mode to be as convenient and uncomplicated as possible, and to be able to interact with people, accurately understand people's interaction intentions, and pass key information to the robotic arm system to perform operations.

[0003] Issuing tasks to robots through language instructions is the most natural way of human-computer interaction for humans. However, single language interaction is difficult to cope with complex daily life scenes. For example, when the surrounding environment is noisy, language interaction may be wrong, or when the language description is vague, it is difficult to accurately locate the target object through language and visual matching. In addition, since the robot's understanding of the task through language and visual fusion lacks interaction with people, if the visual recognition is wrong, it will directly lead to the failure of task execution and reduce the success rate of understanding the interactive intention. The touch screen can visualize the robot's task understanding status. For example, after locating the target object through the matching of language and visual information, the positioning result is visualized on the touch screen interface. In response to the above problems, the present invention proposes a multi-modal interactive fusion intention understanding method, which flexibly combines multiple information such as language, touch screen, and vision to achieve convenient and accurate task communication, and based on this, a rapid learning system is constructed to continuously enhance the robot's interactive understanding ability. Summary of the invention

[0004] After receiving language instructions, the robot understands the key information in the language and matches it with the visual information to obtain the posture information of the object to be operated and guide the execution of the operation.

[0005] In order to improve the success rate of the robot's intention understanding in unstructured and complex daily life scenarios, and ensure the accuracy, generalization and high time efficiency of the robot's understanding of long-sequence multi-task instructions, for the semantic understanding of the robot's operation instructions for picking up and placing objects, the present invention constructs a key information extraction architecture based on a large language model in the task semantic understanding module, including the efficient extraction of multiple key positions and action text information, and performs multi-modal interactive fusion with vision and touch screen to achieve correct understanding of the instruction intention.

[0006] First, based on the construction of a rich human instruction dataset (long sequence multi-task instructions, various tones, scenes, forward and reverse order request instructions), the language processing model is trained through the BERT (Bidirectional Encoder Representation from Transformers) large language model to realize input sentences (which can be long sequence multi-task instructions) and output key action sequences, as shown in formula (1).

[0007] SV=((v1][target 1],(v2)[target 2],…,(v i )[target i]…,(v n )[target n]) (1)

[0008] In the formula, v i represents a text action sequence, i=1,…n, and [target i] represents the target object of interest. Experimental verification shows that the advantages of this model include faster reasoning speed than ChatGpt4.0, better accuracy in understanding the intent of long-sequence multi-task instructions than the method based on grammar analysis and basically the same as ChatGpt4.0.

[0009] The obtained semantic key information matches the visual target detection result, and the key information mentioned in the instruction can be quickly located. In order to confirm the robot's understanding of the key information, the present invention uses PyQt technology to design a touch screen interaction interface, and uses the visual field interface in the touch screen to visualize the target object and related items understood by the robot (at the same time, voice interaction with people "Do you need this item?"). People observe the robot's understanding on the touch screen interface. If it is consistent with the target item required by the person to operate, click "OK" or voice interaction on the touch screen, and the robot completes the picking operation; on the contrary, if the person observes that the operated item displayed on the touch screen interface is wrong, or multiple, or the target item is not identified, the person can click the correct target item on the touch screen. After the robot obtains the target item position correction information, the SegmentAnythingModel (SAM) method is used to accurately extract the correct target item position information in the image and the RGB color information, depth information and point cloud of the corresponding area, which are used for the robot arm to complete the object picking operation. At the same time, the correction information is sent to the fast learning system proposed by the present invention so that the robot can remember the information and provide the correct recognition result when encountering a similar situation next time.

[0010] According to the above technical solution, the present invention has the following advantages:

[0011] The present invention provides an intention understanding method that integrates language, touch screen, and vision multi-mode interactions. The advantages of this method are: 1) the key information extraction model constructed based on a large language model has fast semantic parsing speed and high accuracy; 2) the interactive integration of language, vision, and touch screen enables the robot's understanding of the task to be visualized on the touch screen, and when the language description is vague or the visual positioning is incorrect, the robot can correct the erroneous understanding and memorize it through human-computer touch screen interaction, which effectively improves the robot's recognition ability while increasing the success rate of the robot's task execution. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 A large language model for extracting key information of interaction, a classification table of post-processing actions, and experimental results. (a) The basic architecture of the large language model network for extracting key information of interaction; (b) The conversion categories of various verbs during post-processing; (c.1) Comparison of the long-sequence multi-task instruction capabilities of different models - reasoning speed (unit: ms); (c.2) Comparison of the long-sequence multi-task instruction capabilities of different models - accuracy (unit: %).

[0013] Figure 2 Flow chart of interactive intention understanding based on multimodal information fusion. (a) A1 semantic understanding module and A2 visual recognition module; (b) interactive intention understanding based on the fusion of voice, touch screen and visual information.

[0014] Figure 3 Design of a fast learning system; (a) new knowledge dynamic storage stage; (b) RAG retrieval enhancement generation stage. DETAILED DESCRIPTION

[0015] The present invention is described in detail below with reference to the accompanying drawings and embodiments, but the protection scope of the present invention is not limited to the following embodiments.

[0016] The technical solution adopted by the present invention is an interactive intention understanding and rapid learning system based on multimodal information fusion. Aiming at the semantic understanding of the operation instructions of the robot to pick up and place objects, a key information extraction architecture based on a large language model is constructed in the task semantic understanding module, including the extraction of multiple key positions and action text information, and the extracted information is fused with vision and touch screen for multimodal interaction to achieve correct understanding of the instruction intention.

[0017] The key information extraction architecture based on the large language model is based on the BERT large language model, and the language processing model is trained by the BERT large language model; the language processing model includes a human instruction data set of long-sequence multi-task instructions and various tones, scenes, forward and reverse request instructions, so as to input long-sequence multi-task instruction sentences and output key action sequences; as shown in formula (1).

[0018] SV=((v1)[target 1],(v2)[target 2],…,(v i )[target i]…,(v n )[targetn]) (1)

[0019] In the formula, v i Represents a text action sequence, i=1,…n, and [target i] represents the target object of interest.

[0020] Furthermore, in the key information extraction architecture based on the large language model, the obtained semantic key information is matched with the visual target detection results to quickly locate the key information mentioned in the instruction. In order to confirm the robot's understanding of the key information, the touch screen interaction interface is designed using PyQt technology, and the target object and related items understood by the robot are visualized using the visual field interface in the touch screen. The robot's understanding is observed on the touch screen interface. If it is consistent with the target object required to be operated, click "OK" on the touch screen or voice interaction, and the robot completes the picking operation; on the contrary, if it is observed that the operated object displayed on the touch screen interface is wrong, or multiple, or the target object is not recognized, click the correct target object on the touch screen. When the robot obtains the target object position correction information, the SAM method is used to accurately extract the correct target object position information in the image and the RGB color information, depth information and point cloud of the corresponding area, which are used for the robotic arm to complete the picking operation. The correction information is sent to the proposed fast learning system so that the robot can remember the information.

[0021] Furthermore, the BERT-based large language model is a BERT+Bi-LSTM+CRF architecture, where the BERT model is a pre-trained RoBERTa for embedding and encoding. The output of the BERT model is sent to the Bi-LSTM layer to better capture the bidirectional dependencies of the sentence. Finally, a CRF linear layer is added (as shown in formula (2)) to map the output of the Bi-LSTM layer to the label space, and the output is sent to the CRF linear layer.

[0022] O t =Wh t +b (2)

[0023] In formula (2), W and b are the weight matrix and bias vector of the CRF linear layer respectively. t h is the output of the CRF linear layer and will be transferred to the CRF linear layer. t is the output of the Bi-LSTM layer.

[0024] Decoding and loss calculation are performed at the CRF layer. The last layer based on the BERT large language model is the post-processing layer, which is a newly added part. The output of the CRF linear layer is regularized to transform it into an action sequence that can be executed by the robot, including the following processing: (1) connecting characters with the same entity label; (2) arranging the entity labels in the order of first position, first action, second position, second action, ..., xth position, xth action; (3) optimizing the task executableness and classifying the various actions into three categories: "pick up, place, and other".

[0025] Furthermore, the input based on the BERT large language model is an interactive intention understanding system based on multimodal information fusion, including a semantic understanding module, a visual recognition module, and a fast learning module.

[0026] In the semantic understanding module, when a person issues a voice command, the voice is converted into a text string through the Open WhisperAPI model and input into the proposed large language model to extract key interaction information, which lists the operation actions and target operation objects or positions in order.

[0027] In the visual recognition module, the visual information obtained from the depth camera is displayed in the form of a video stream on the touch screen interface, and is input into the visual recognition model to detect the category and location information of the objects in the scene. The SAM method is used to segment the scene objects in the image. SAM is a pre-trained image segmentation model that can perform zero-sample transfer to new images and tasks without additional training. The image is input into SAM to obtain the information of the object segmentation in the scene. Then, the category name is added to each object segmentation area in combination with the ChatGPT prompt word. For occasional segmentation errors, a category and segmentation consistency correction submodule is added. When the adjacent segment blocks in the scene have two or more different IDs and the category name is consistent, the segment blocks are merged into one block and described with the category name. The visual recognition module finally outputs reliable object category names and segmentation location information in the scene. The real-time open vocabulary object detection model composed of SAM segmentation, ChatGPT and correction submodules is applied in scenes with messy and diverse objects such as daily homes to obtain the category names and location information of objects in the scene.

[0028] After obtaining the sequence of key interactive actions and related operation items or positions through the semantic understanding module, the information is matched with the output of the visual recognition module. If all the related items or positions are matched, the position information of the related items in the image is obtained from the output of the visual recognition module and transmitted to the touch screen. The interactive items and other related item position information are visually displayed in the scene video stream window of the touch screen interface. At the same time, voice interaction prompts are given. If you think the robot's understanding is correct, click the "OK" button on the touch screen, or reply "yes" by voice. After obtaining the interactive confirmation information, the robot sends the visual information of the related items to the robot's operation planning and execution module, so that the robot can complete the task of picking up or putting something somewhere.

[0029] If the user thinks the robot's understanding is wrong or the match is unsuccessful, the robot will prompt this through voice interaction. The user clicks on the target interactive object on the touch screen video stream, searches for the segmentation block where this coordinate position is located in the SAM segmentation result, and displays this segmentation block visually on the touch screen. At the same time, the user interacts with the user through touch screen or voice to confirm whether the robot's understanding is correct. If the robot understands correctly, the relevant item information is sent to the robot planning and execution module; otherwise, the process of clicking the image on the touch screen is repeated.

[0030] If the surrounding environment is noisy or the target object is cluttered, start directly by clicking the video stream on the touch screen, segment and identify the target object, and perform touch-screen interactive confirmation to obtain the target object's location information.

[0031] The rapid learning module includes: (a) new knowledge dynamic storage stage and (b) RAG retrieval enhancement generation stage.

[0032] In the RAG retrieval enhancement generation phase, the internal knowledge of the large language model is dynamically and collaboratively merged with the external database to enhance the recognition ability of the large language model.

[0033] In the new knowledge dynamic storage stage, when the robot misunderstands the category name of an item in the touch screen visualization display, it quickly adds the category name of the new item and its visual attribute features to the dynamic knowledge graph library by clicking on the item on the touch screen to obtain its corresponding image segmentation block and inputting the correct category name by voice. The knowledge vector representation TransE model is used to embed entities, relationships and attributes into a low-dimensional vector space to achieve efficient representation of knowledge in the graph and complete dynamic storage of new knowledge.

[0034] When receiving the user's instruction, a similarity comparison is performed in the knowledge base in the form of a vector to ensure that the retrieved information is similar enough. In the generation phase, the original text is combined with the retrieved visual attribute features as input and sent to the large language model. The large language model gives the category name of the items similar to the visual attribute features in the current scene image while comprehensively considering the input content.

[0035] Example

[0036] Attached Figure 1 In the figure, (a) is the basic architecture of the large language model network for extracting key interactive information. The BERT+Bi-LSTM+CRF architecture is adopted (Bi-LSTM is the abbreviation of Bi-directional Long Short-Term Memory, and CRF is the abbreviation of Conditional Random Fields), in which the BERT model uses the pre-trained RoBERTa for embedding and encoding. The output of the BERT model is transmitted to the Bi-LSTM layer to better capture the bidirectional dependencies of the sentence and improve the expression and generalization capabilities of the large language model. Finally, a CRF linear layer is added to map the output of the Bi-LSTM layer to the label space, and the output is transmitted to the CRF linear layer.

[0037] O t =Wh t +b

[0038] In formula (2), W and b are the weight matrix and bias vector of the CRF linear layer respectively. t h is the output of the CRF linear layer and will be transferred to the CRF linear layer. t is the output of the Bi-LSTM layer.

[0039] The CRF layer mainly performs decoding and loss calculation. The last layer of the large language model is the post-processing layer. As a new part, the output of the CRF linear layer is regularized to transform it into an action sequence that can be executed by the robot. The main processing includes the following: (1) connecting characters with the same entity label; (2) arranging the entity labels in the order of first position, first action, second position, second action, ...; (3) optimizing the task executableness and classifying the various actions into three categories: "pick up, place, and other" (see Appendix). Figure 1 (c.1) shows in detail the transformation categories of various verbs during post-processing).

[0040] Attached Figure 1In the figure, Figure (c.1) and Figure (c.2) compare the processing time and semantic understanding accuracy of three semantic understanding models under different numbers of basic tasks (each model was tested 20 times for each number of basic tasks). SpaCy is a natural language processing library that can be used to efficiently process natural language tasks. The other two are the large language model ChatGpt4.0 and the method proposed in the present invention. The method in the present invention is significantly better than ChatGpt4.0 in terms of reasoning speed, and has no obvious speed disadvantage compared with the dependency grammar analysis method (SpaCy). However, as the number of basic tasks increases in accuracy, the performance of the SpaCy model decreases significantly, and even becomes unusable, while the accuracy of ChatGpt4.0 decreases slightly. The accuracy of the method proposed in the present invention is similar to that of ChatGpt4.0, and is slightly higher than ChatGpt4.0 when the number of basic tasks increases.

[0041] Attached Figure 2 This is a flowchart of the interactive intention understanding process based on multi-modal information fusion, where A1 represents the semantic understanding module, A2 represents the visual recognition module, and A3 represents the fast learning module. Figure 2 As shown in (a), in A1, after a person issues a voice command, the voice is converted into a text string through the Open WhisperAPI model and input into the proposed large language model to extract key interaction information, which lists the operation actions and target operation objects or positions in order.

[0042] In the A2 visual recognition module, the visual information obtained from the depth camera displays the scene information in the form of a video stream on the touch screen interface, and is input into the visual recognition model on the other hand to detect the category and location information of the objects in the scene. The present invention uses the SegmentAnything Model (SAM) method to segment the scene objects in the image. SAM is a pre-trained image segmentation model that can perform zero-sample migration on new images and tasks without additional training. By inputting the image into the SAM model, the information of the segmentation of objects in the scene can be obtained. Then, combined with the ChatGPT prompt words, such as "I have labeled a bright numeric ID for each visual object in the image. Please numerate their names.", the category name is further added to each object segmentation area. In response to occasional segmentation errors, the present invention adds a category and segmentation consistency correction submodule. When the adjacent segmentation blocks in the scene have two or more different IDs, but the category names are consistent, the segmentation blocks are merged into one block and described with the category name. As a result, the A2 visual recognition module can finally output reliable item category names and segmentation location information in the scene (in the robot's field of view). The real-time open-vocabulary object detection model composed of SAM segmentation, GPT and correction modules is suitable for use in scenes with cluttered and diverse objects such as daily homes to obtain the category names and location information of objects in the scene.

[0043] As attached Figure 2 As shown in (b), after obtaining the sequence of interactive key actions and related operation items or positions through the A1 module, the information is matched with the output of the A2 module (the name and position of the item category in the robot's field of view). If all the related items or positions are matched, the position information of the related items in the image is obtained from the output of the A2 module and transmitted to the touch screen. The interactive items and other related item position information are visually displayed in the scene video stream window of the touch screen interface (for example, by highlighting the bold enclosing box, etc.). At the same time, the voice interaction prompts, "Is this the item you need?" If the person thinks that the robot's understanding is correct, then he can click the "OK" button on the touch screen, or reply "Yes" by voice. After obtaining the interactive confirmation information, the robot sends the visual information of the related items (image, depth and point cloud information) to the robot's operation planning and execution module, so that the robot can complete the task of picking up or placing something somewhere.

[0044] If the user believes that the robot's understanding is wrong (caused by visual recognition errors) or the match is unsuccessful (such as when the language text description is vague, the matching result is not unique, etc.), the robot can prompt this situation through voice interaction. The user clicks on the target interactive object on the touch screen video stream, searches for the segmentation block where this coordinate position is located in the SAM segmentation result, and displays this segmentation block visually on the touch screen. At the same time, the user interacts with the user through touch screen or voice to confirm whether the robot's understanding is correct. If the robot understands correctly, the relevant item information is sent to the robot planning and execution module; otherwise, the above touch screen clicking image process is repeated.

[0045] If the surrounding environment is noisy or the target object is cluttered, you can also start directly by clicking on the video stream on the touch screen, segmenting and identifying the target object, and performing touch-screen interactive confirmation to obtain the target object's location information.

[0046] Attached Figure 3 It is a fast learning system design, including two stages: (a) the dynamic storage stage of new knowledge and (b) the RAG (Retrieval-Augmented Generation) retrieval enhancement generation stage. Due to the fact that new objects often appear in an open home environment, the robot is required to have continuous and rapid learning capabilities and to be able to continuously update its visual recognition capabilities. In response to this problem, the present invention constructs a RAG retrieval enhancement generation system to enhance the recognition capabilities of the large language model by dynamically and collaboratively merging the intrinsic knowledge of the large language model with an external database. When the user finds that the robot misunderstands the category name of an item in the touch-screen visualization display, the user can quickly add the category name of the new item and its visual attribute features to the dynamic knowledge graph library by clicking on the item on the touch screen to obtain its corresponding image segmentation block and voice inputting the correct category name. The knowledge vector representation TransE model is used to embed entities, relationships and attributes into a low-dimensional vector space to achieve efficient representation of knowledge in the graph (see Appendix). Figure 3 (a)).

[0047] Attached Figure 3In (b), when the model receives the user's instruction, it uses a vector to perform a similarity comparison in the knowledge base to ensure that the retrieved information is similar enough. In the generation stage, the original text is combined with the retrieved visual attribute features as input and sent to the language model. The final model gives the category name of objects with similar visual attribute features in the current scene image while comprehensively considering the input content. The advantage of the RAG system is that it can use a large amount of external knowledge for reasoning and judgment, and is suitable for handling situations where knowledge is frequently updated. In the present invention, although the A2 visual recognition module can recognize open vocabulary objects, there may be recognition errors. By combining touch screen input and applying the A3 fast learning system, the ability of the A2 visual recognition module can be significantly and continuously enhanced, which is one of the important ways to improve the capabilities of robots.

Claims

1. An interactive intention understanding and rapid learning system based on multi-modal information fusion, which is aimed at the semantic understanding of the operation instructions of robots to pick up and place objects; it is characterized by: In the task semantic understanding module, a key information extraction architecture based on a large language model is built, including the extraction of multiple key positions and action text information, and the extracted information is integrated with vision and touch screen for multi-modal interaction to achieve the correct understanding of the instruction intent; The key information extraction architecture based on the large language model is based on the BERT large language model, and the language processing model is trained by the BERT large language model; the language processing model includes a human instruction data set of long-sequence multi-task instructions and various tones, scenes, forward and reverse request instructions, so as to input long-sequence multi-task instruction sentences and output key action sequences; as shown in formula (1); SV=((v1)[target 1],(v2)[target 2],…,(v i )[target i]…,(v n )[targetn]) (1) In the formula, v i Represents a text action sequence, i=1,…n, and [target i] represents the target object of interest.

2. The interactive intention understanding and rapid learning system based on multimodal information fusion according to claim 1 is characterized in that: In the key information extraction architecture based on the large language model, the semantic key information obtained is matched with the visual target detection results to quickly locate the key information mentioned in the instruction; in order to confirm the robot's understanding of the key information, the touch screen interaction interface is designed using PyQt technology, and the target object and related items understood by the robot are visualized using the visual field interface in the touch screen; the robot's understanding is observed on the touch screen interface, and if it is consistent with the target object required to be operated, "OK" is clicked on the touch screen or voice interaction is performed, and the robot completes the picking operation; on the contrary, if it is observed that the operated object displayed on the touch screen interface is wrong, or multiple, or the target object is not recognized, the correct target object is clicked on the touch screen; when the robot obtains the target object position correction information, the SAM method is used to accurately extract the correct target object position information in the image and the RGB color information, depth information and point cloud of the corresponding area, which are used for the robotic arm to complete the picking operation; the correction information is sent to the proposed fast learning system so that the robot can remember the information.

3. The interactive intention understanding and rapid learning system based on multimodal information fusion according to claim 1 is characterized in that: The BERT-based large language model is a BERT+Bi-LSTM+CRF architecture, where the BERT model is a pre-trained RoBERTa for embedding and encoding; the output of the BERT model is sent to the Bi-LSTM layer to better capture the bidirectional dependencies of the sentence; finally, a CRF linear layer is added (as shown in formula (2)) to map the output of the Bi-LSTM layer to the label space, and the output is sent to the CRF linear layer; O t =Wh t +b (2) In formula (2), W and b are the weight matrix and bias vector of the CRF linear layer respectively; t is the output of the CRF linear layer and will be transmitted to the CRF linear layer; h t is the output of the Bi-LSTM layer; Decoding and loss calculation are performed at the CRF layer. The last layer based on the BERT large language model is the post-processing layer, which is a newly added part. The output of the CRF linear layer is regularized to transform it into an action sequence that can be executed by the robot, including the following processing: (1) connecting characters with the same entity label; (2) arranging the entity labels in the order of first position, first action, second position, second action, ..., xth position, xth action; (3) optimizing the task executableness and classifying the various actions into three categories: "pick up, place, and other".

4. The interactive intention understanding and rapid learning system based on multimodal information fusion according to claim 3 is characterized in that: The input of the BERT large language model is an interactive intention understanding system based on multimodal information fusion, including a semantic understanding module and a visual recognition module; The sequence of key interactive actions and related operation items or positions is obtained through the semantic understanding module, and the information is matched with the output of the visual recognition module; if all related items or positions are matched, the position information of the related items in the image is obtained from the output of the visual recognition module, and it is transmitted to the touch screen, and the interactive items and other related items are visually displayed in the scene video stream window of the touch screen interface. At the same time, voice interaction prompts are given; if it is believed that the robot's understanding is correct, click the "OK" button on the touch screen, or reply "Yes" by voice; after obtaining the interactive confirmation information, the robot sends the visual information of the related items to the robot's operation planning and execution module, so that the robot can complete the task of picking up or putting something somewhere; If the user thinks that the robot's understanding is wrong or the match is unsuccessful, the robot will prompt this situation through voice interaction; the user clicks the target interactive object on the touch screen video stream, searches for the segmentation block where the coordinate position is located in the SAM segmentation result, and displays the segmentation block visually on the touch screen, and interacts with the user through touch screen or voice to confirm whether the robot's understanding is correct; if the robot's understanding is correct, the relevant object information is sent to the robot planning and execution module; otherwise, the process of touching the screen and clicking the image is repeated; If the surrounding environment is noisy or the target object is cluttered, start directly by clicking the video stream on the touch screen, segment and identify the target object, and perform touch-screen interactive confirmation to obtain the target object's location information.

5. The interactive intention understanding and rapid learning system based on multimodal information fusion according to claim 4 is characterized in that: In the semantic understanding module, when a person issues a voice command, the voice is converted into a text string through the Open Whisper API model and input into the proposed large language model to extract key interaction information, which lists the operation actions and target operation objects or positions in order.

6. The interactive intention understanding and rapid learning system based on multimodal information fusion according to claim 4 is characterized in that: In the visual recognition module, the visual information obtained from the depth camera is displayed on the touch screen interface in the form of a video stream, and is input into the visual recognition model to detect the category and location information of objects in the scene. The SAM method is used to segment the scene objects in the image. SAM is a pre-trained image segmentation model that can perform zero-shot transfer to new images and tasks without additional training. The image is fed into SAM to obtain information about the object segmentation in the scene. Then, the ChatGPT prompts are used to add category names to each object segmentation region. In response to occasional segmentation errors, a category and segmentation consistency correction submodule is added. When adjacent segments in the scene have two or more different IDs and the category name is the same, the segments are merged into one and described by the category name; the visual recognition module finally outputs reliable object category names and segmentation location information in the scene; the real-time open vocabulary object detection model composed of SAM segmentation, ChatGPT and correction submodules is applied in daily home scenes with cluttered and diverse objects to obtain the category names and location information of objects in the scene.

7. The interactive intention understanding and rapid learning system based on multimodal information fusion according to claim 4 is characterized in that: The interactive intention understanding system based on multimodal information fusion also includes a fast learning module, including: (a) new knowledge dynamic storage stage and (b) RAG retrieval enhancement generation stage; In the RAG retrieval enhancement generation phase, the internal knowledge of the large language model is dynamically and collaboratively merged with the external database to enhance the recognition ability of the large language model; In the new knowledge dynamic storage stage, when the robot is found to have misunderstood the category name of an item in the touch screen visualization display, it can quickly add the category name of the new item and its visual attribute features to the dynamic knowledge graph library by clicking on the item on the touch screen to obtain its corresponding image segmentation block and input the correct category name by voice. The knowledge vector representation TransE model is used to embed entities, relationships and attributes into a low-dimensional vector space to achieve efficient representation of knowledge in the graph and complete dynamic storage of new knowledge. When the user's instructions are received, a similarity comparison is performed in the knowledge base in the form of a vector to ensure that the retrieved information is similar enough; in the generation stage, the original text is combined with the retrieved visual attribute features as input and sent to the large language model; the large language model, taking into account the input content, gives the category name of the items with similar visual attribute features in the current scene image.

Citation Information

Patent Citations

  • Accompany robot and control method thereof

    CN109545195A

  • Structured information extraction method and device based on multi-element labeling strategy

    CN113836891A

  • Human-computer interaction intention understanding method based on visual language multi-modal fusion

    CN117725554A

Cited By

  • Task processing method and device for robot, electronic equipment and robot

    CN120116232A

  • Man-machine interaction method and man-machine interaction system based on large model assistance

    CN120199245A

  • Voice intention recognition method, electronic equipment and storage medium

    CN120279912A

  • Language control interaction method and system based on AR and VR

    CN120913568A

  • Task planning method and system for robot

    CN121061907A