Scene perception and interaction method and system based on multi-modal large model

The scene semantic information is obtained through the multimodal big model, combined with preset interaction conditions, the robot's perception and interaction problems in the dynamic environment are solved, and efficient intelligent interaction and rapid response are achieved.

CN120347789APending Publication Date: 2025-07-22SHANDONG NEW GENERATION INFORMATION IND TECH RES INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510363859.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing technology lacks autonomous perception and interaction capabilities of robots in dynamic and complex environments, lacks a multimodal data fusion framework and dynamic interaction decision-making mechanism, and is difficult to adapt to the needs of complex scenarios.

Method used

A multimodal large model is used to collect scene pictures through visual sensors, obtain scene semantic information, and combine preset interaction conditions to perform perception and interaction judgment, and perform corresponding behaviors.

Benefits of technology

The robot's perception accuracy and interaction efficiency in dynamic environments have been improved, and the intelligent interaction capability and response speed have been significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120347789A_ABST
    Figure CN120347789A_ABST
Patent Text Reader

Abstract

The invention discloses a scene perception and interaction method and system based on a multi-modal large model, and belongs to the technical field of robots, and the method comprises the following steps: collecting picture information in a scene through a visual sensor; inputting the picture information into a multi-modal large model to obtain scene semantic information; sensing a scene based on the scene semantic information; comparing the semantic information of the current scene with a preset interaction condition, and judging whether environment interaction needs to be carried out or not; and when the interaction condition is met, executing a corresponding interaction behavior. According to the method, the limitation of single-mode sensing is solved, the intelligent interaction capability of the robot is improved, and the response speed of the robot in a dynamic environment is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of robotics, and more specifically, to a method and system for scene perception and interaction based on a multimodal large model. Background Art

[0002] With the rapid development of robotics, the autonomous perception and interaction capabilities of robots in dynamic and complex environments have become a research hotspot. With the breakthrough progress of multimodal large models, their powerful cross-modal data understanding capabilities provide new solutions for robot perception and interaction technologies. However, most existing studies remain at the single-task verification in laboratory environments and have not yet formed a multimodal data fusion framework and dynamic interaction decision-making mechanism applicable to real scenarios. In addition, how to combine the semantic understanding capabilities of large models with the real-time control requirements of robots to achieve efficient scene perception and interaction collaboration remains a technical problem to be solved urgently.

[0003] Existing technologies mostly rely on pre-programmed fixed rules or simple trigger conditions, lacking dynamic understanding of the environment and intelligent interaction capabilities. For example, robots cannot autonomously judge whether they need to interact with the environment based on scene semantic information, or their interaction behaviors are single and difficult to meet the requirements of complex scenarios. Summary of the Invention

[0004] The technical task of the present invention is to address the above deficiencies by providing a method and system for scene perception and interaction based on a multimodal large model, which solves the limitations of single-modal perception, enhances the intelligent interaction capabilities of robots, and significantly improves the response speed of robots in dynamic environments.

[0005] The technical solution adopted by the present invention to solve its technical problems is as follows:

[0006] A method for scene perception and interaction based on a multimodal large model, the implementation of which includes the following steps:

[0007] 1) Collect picture information in the scene through a visual sensor;

[0008] 2) Input the picture information into the multimodal large model to obtain scene semantic information;

[0009] 3) Perceive the scene based on the scene semantic information;

[0010] 4) Compare the current scene semantic information with preset interaction conditions to determine whether environmental interaction is required;

[0011] 5) When the interaction conditions are met, perform corresponding interaction behaviors.

[0012] This method significantly improves the perception accuracy and interaction efficiency of the robot in a dynamic and complex environment through multi-modal data fusion and intelligent interaction determination.

[0013] Further, the multi-modal large model includes a multi-modal feature extraction model with an input of pictures and an output of feature vectors.

[0014] The multi-modal large model includes a multi-modal large language model with inputs of pictures and semantic information prompts (Prompts).

[0015] Further, the multi-modal feature extraction model with an input of pictures and an output of feature vectors includes a CLIP model.

[0016] The multi-modal large language model with inputs of pictures and semantic information prompts (Prompts) includes methods such as LLAVA and miniCPM.

[0017] The semantic information prompt Prompt is a description of what type of information to obtain from the picture, including a description prompt for general information and a description prompt for specific information. The description prompt for general information is a prompt for extracting information in all scenarios. For example, "Please describe the main content in the picture." is the general information to be extracted from all pictures. The description prompt for specific information is a prompt formulated for a certain special perception requirement and interaction requirement. For example, "Check whether the screen is in the off state?" is a prompt for specific information when the requirement contains the determination of the screen state.

[0018] Further, the scene semantic information includes the feature vectors extracted by the multi-modal feature extraction model and the image descriptions extracted by using the multi-modal large language model.

[0019] Further, based on the scene semantic information, the scene is perceived, including:

[0020] The detection and recording of static objects in the environment, including real-time monitoring and recording of object categories, object positions, environmental states, and scene changes.

[0021] The detection and recording of dynamic objects in the environment, including the monitoring and recording of personnel categories, the number of personnel, personnel behaviors, and the operating states of machines.

[0022] Further, the interaction conditions include abnormal interaction conditions and normal interaction conditions.

[0023] The abnormal interaction conditions include static object abnormalities, including object abnormalities and environmental abnormalities; and also include dynamic object abnormalities, including personnel abnormalities and machine abnormalities.

[0024] Normal interaction-based interaction conditions, including the reporting conditions for the semantic information of the current environment, the conditions for personnel to await interaction when there are people in the environment, the conditions for reporting the status of objects in the environment, etc.;

[0025] The said step 4),

[0026] When abnormal interaction conditions are met, the interaction behaviors carried out by the abnormal interaction conditions include the reporting and reminder of abnormal states within the robot system, the real-time reporting of abnormalities, the voice reminder and intervention for personnel abnormalities, the reporting and robotic arm operation intervention for machine abnormalities, and the robotic arm operation intervention for abnormalities in the environmental state, etc.;

[0027] The interaction behaviors carried out when normal interaction conditions are met include the reporting and reminder of the semantic information of the current scene within the robot system, the initiation of interaction invitations for personnel in the environment through voice, and the evaluation and status reporting based on the objects in the environment.

[0028] Further, for the said step 5), the interaction behaviors include actions such as grasping and moving objects in the environment based on the interaction content and navigation planning.

[0029] The present invention also claims protection for a scene perception and interaction system based on a multimodal large model, including:

[0030] A visual sensor for collecting picture information in the scene;

[0031] A multimodal large model for processing the said picture information and extracting scene semantic information;

[0032] A perception module for perceiving the scene based on the scene semantic information;

[0033] An interaction determination module for comparing the current scene semantic information with preset interaction conditions;

[0034] An interaction execution module for executing corresponding interaction behaviors when the interaction conditions are met;

[0035] The said multimodal large model supports real-time calculation to ensure the rapid response of the robot in a dynamic environment;

[0036] The said perception module can dynamically record and retrieve the objects, environment, and status in the scene;

[0037] The said interaction execution module supports various interaction behaviors, including voice reporting, visual grasping, and navigation planning;

[0038] This system can implement the above-mentioned scene perception and interaction method based on a multimodal large model.

[0039] The present invention also claims protection for a scene perception and interaction device based on a multimodal large model, including: at least one memory and at least one processor;

[0040] The at least one memory is used for storing machine-readable programs;

[0041] The at least one processor is used for calling the machine-readable programs to implement the above method.

[0042] The present invention also claims protection for a computer-readable medium, on which computer instructions are stored, and when the computer instructions are executed by a processor, the above method is implemented.

[0043] Compared with the prior art, a scene perception and interaction method and system based on a multimodal large model of the present invention have the following beneficial effects:

[0044] The present invention inputs the scene picture information collected by a visual sensor into a multimodal large model to obtain rich scene semantic information, including objects, environments, states, etc., and based on this information, realizes comprehensive scene perception and intelligent interaction; through the efficient data processing ability of the multimodal large model, the limitation of single-modal perception is solved; through the dynamic matching of preset interaction conditions and scene semantic information, the intelligent interaction ability of the robot is improved; at the same time, by utilizing the real-time computing advantage of the multimodal large model, the response speed of the robot in a dynamic environment is significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 is a flowchart of a scene perception and interaction method based on a multimodal large model provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] The following further describes the present invention with reference to specific embodiments.

[0047] The technical solution adopted by the embodiment of the present invention to solve its technical problems is:

[0048] A scene perception and interaction method based on a multimodal large model, and the implementation of this method includes the following steps:

[0049] 1. Collect picture information in the scene through a visual sensor;

[0050] 2. Input the picture information into a multimodal large model to obtain scene semantic information;

[0051] 3. Perceive the scene based on the scene semantic information;

[0052] 4. Compare the current scene semantic information with preset interaction conditions to determine whether environmental interaction is required;

[0053] 5. When the interaction conditions are met, perform the corresponding interaction behaviors.

[0054] Among them, the multimodal large model includes a multimodal feature extraction model with an input of pictures and an output of feature vectors, including but not limited to the CLIP model.

[0055] The multimodal large model includes a multimodal large language model with an input of pictures and semantic information prompts (Prompts), including but not limited to methods such as LLAVA and miniCPM.

[0056] The semantic information prompt Prompt is a description of what type of information is to be obtained from the picture, including a description prompt for general information and a description prompt for specific information. The general type of prompt is a prompt for extracting information in all scenarios. For example, "Please describe the main content in the picture." is the general information to be extracted from all pictures. The description prompt for specific information is a prompt formulated for a certain special perception requirement and interaction requirement. For example, "Check whether the screen is in the off state?" is a prompt for specific information when the requirement contains the determination of the screen state.

[0057] The scene semantic information includes the feature vectors extracted by the multimodal feature extraction model and the image descriptions extracted by using the multimodal large language model.

[0058] Based on the scene semantic information, perceive the scene; including:

[0059] Detect and record static objects in the environment, including real-time monitoring and recording of object categories, object positions, environmental states, and scene changes;

[0060] Detect and record dynamic objects in the environment, including monitoring and recording of personnel categories, the number of personnel, personnel behaviors, and the operating states of machines.

[0061] The interaction conditions include abnormal interaction conditions and normal interaction conditions:

[0062] The abnormal interaction conditions include static object abnormalities, including object abnormalities and environmental abnormalities; and also include dynamic object abnormalities, including personnel abnormalities and machine abnormalities.

[0063] The normal interaction conditions include the reporting conditions for the current environmental semantic information, the personnel interaction conditions when there are people in the environment, the conditions for reporting the object states in the environment, etc.

[0064] Step 4,

[0065] When the abnormal interaction conditions are met, the interaction behaviors carried out by the abnormal interaction conditions include the reporting and reminder of the abnormal state within the robot system, the real-time broadcast of the abnormality, the voice reminder intervention for personnel abnormalities, the broadcast and robotic arm operation intervention for machine abnormalities, and the robotic arm operation intervention for abnormalities in the environmental state, etc.;

[0066] The interaction behaviors carried out when the normal interaction conditions are met include the reporting and reminder of the current scene semantic information within the robot system, the voice-initiated interaction invitation for the people in the environment, and the evaluation and status broadcast based on the objects in the environment.

[0067] In step 5, the interaction behaviors include actions such as the grasping and moving of objects in the environment and the navigation planning based on the interaction content.

[0068] Through this method, during the driving process of the robot, the visual sensor collects the picture information in the scene and inputs it into the multi-modal large model to obtain the scene semantic information, including the perception information such as objects, environment, and status. Based on the scene semantic information, the robot can perceive the scene, realize object recognition, environmental status recording, retrieval, and feedback. At the same time, by comparing the current scene semantic information with the preset interaction conditions, it is determined whether to interact with the environment; when the interaction conditions are met, the robot executes corresponding interaction behaviors such as voice broadcast and visual grasping. This method significantly improves the perception accuracy and interaction efficiency of the robot in a dynamic and complex environment through multi-modal data fusion and intelligent interaction determination, has a wide range of application prospects, is applied to the mapping and positioning scenario, in the robot product, and has strong technical advantages in the industry.

[0069] The embodiment of the present invention also provides a scene perception and interaction system based on a multi-modal large model, including:

[0070] A visual sensor for collecting the picture information in the scene;

[0071] A multi-modal large model for processing the picture information and extracting the scene semantic information;

[0072] A perception module for perceiving the scene based on the scene semantic information;

[0073] An interaction determination module for comparing the current scene semantic information with the preset interaction conditions;

[0074] An interaction execution module for executing corresponding interaction behaviors when the interaction conditions are met.

[0075] The multi-modal large model supports real-time calculation to ensure the rapid response of the robot in a dynamic environment.

[0076] The perception module can dynamically record and retrieve objects, environments, and states in the scene.

[0077] The interaction execution module supports multiple interaction behaviors, including voice broadcast, visual grasping, and navigation planning.

[0078] The system can implement the scene perception and interaction method based on the multimodal large model described in the above embodiments. The specific implementation process is as follows:

[0079] 1. Collect picture information in the scene through a visual sensor;

[0080] 2. Input the picture information into the multimodal large model to obtain scene semantic information;

[0081] 3. Perceive the scene based on the scene semantic information;

[0082] 4. Compare the current scene semantic information with the preset interaction conditions to determine whether environmental interaction is required;

[0083] 5. When the interaction conditions are met, execute the corresponding interaction behavior.

[0084] Among them, the multimodal large model includes a multimodal feature extraction model with pictures as input and feature vectors as output, including but not limited to the CLIP model.

[0085] The multimodal large model includes a multimodal large language model with pictures and semantic information prompts (Prompt) as input, including but not limited to methods such as LLAVA and miniCPM.

[0086] The semantic information prompt Prompt is a description of what type of information is to be obtained from the picture, including description prompts for general information and description prompts for specific information. General type prompts are prompts for extracting information in all scenes. For example, "Please describe the main content in the picture." is the general information to be extracted from all pictures. The description prompt for specific information is a prompt formulated for a certain special perception requirement and interaction requirement. For example, "Check whether the screen is in the off state?" is a prompt for specific information when the requirement contains the determination of the screen state.

[0087] The scene semantic information includes the feature vectors extracted by the multimodal feature extraction model and the image descriptions extracted by the multimodal large language model.

[0088] The perceiving the scene based on the scene semantic information includes:

[0089] Detecting and recording static objects in the environment, including real-time monitoring and recording of object categories, object positions, environmental states, and scene changes;

[0090] Detection and recording of dynamic objects in the environment, including monitoring and recording of personnel categories, the number of personnel, personnel behavior, and the operating status of machines.

[0091] The interaction conditions include abnormal interaction conditions and normal interaction conditions:

[0092] Abnormal interaction conditions include static object abnormalities, including object abnormalities and environmental abnormalities; they also include dynamic object abnormalities, including personnel abnormalities and machine abnormalities.

[0093] Normal interaction conditions include conditions for reporting semantic information of the current environment, conditions for personnel to await interaction when there are people in the environment, conditions for reporting the status of objects in the environment, etc.

[0094] The said step 4,

[0095] When the abnormal interaction conditions are met, the interaction behaviors carried out by the abnormal interaction conditions include reporting and reminding of abnormal states within the robot system, real-time reporting of abnormalities, voice reminder and intervention for personnel abnormalities, reporting and robotic arm operation intervention for machine abnormalities, and robotic arm operation intervention for abnormalities in the environmental state, etc.;

[0096] The interaction behaviors carried out when the normal interaction conditions are met include reporting and reminding of the semantic information of the current scene within the robot system, initiating an interaction invitation by voice for the people in the environment, and evaluating and reporting the status based on the objects in the environment.

[0097] The said step 5, the interaction behaviors include actions such as grasping, moving, and navigation planning of objects in the environment based on the interaction content.

[0098] An embodiment of the present invention also provides a scene perception and interaction device based on a multi-modal large model, including: at least one memory and at least one processor;

[0099] The said at least one memory is used to store machine-readable programs;

[0100] The said at least one processor is used to call the machine-readable program to implement the scene perception and interaction method based on the multi-modal large model described in the above embodiment.

[0101] An embodiment of the present invention also provides a computer-readable medium, on which computer instructions are stored. When the computer instructions are executed by a processor, the method for scenario perception and interaction based on a multimodal large model described in the above embodiments is implemented. Specifically, a system or device equipped with a storage medium can be provided, on which software program code for implementing the functions of any one of the above embodiments is stored, and the computer (or CPU or MPU) of the system or device is caused to read and execute the program code stored in the storage medium.

[0102] In this case, the program code read from the storage medium itself can implement the functions of any one of the above embodiments. Therefore, the program code and the storage medium storing the program code constitute a part of the present invention.

[0103] Examples of the storage medium for providing the program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, the program code can be downloaded from a server computer via a communication network.

[0104] Furthermore, it should be clear that not only can the functions of any one of the above embodiments be implemented by executing the program code read by the computer, but also by causing an operating system or the like operating on the computer based on the instructions of the program code to complete part or all of the actual operations.

[0105] In addition, it can be understood that the program code read from the storage medium is written into the memory provided in the expansion board inserted into the computer or into the memory provided in the expansion unit connected to the computer, and then based on the instructions of the program code, the CPU or the like installed on the expansion board or the expansion unit is caused to execute part and all of the actual operations, thereby implementing the functions of any one of the above embodiments.

[0106] The present invention has been described in detail above with reference to the accompanying drawings and preferred embodiments. However, the present invention is not limited to these disclosed embodiments. Based on the above-mentioned multiple embodiments, those skilled in the art can know that more embodiments of the present invention can be obtained by combining the code review means in the above different embodiments, and these embodiments are also within the protection scope of the present invention.

Claims

1. A scene perception and interaction method based on a multimodal large model, characterized in that, The implementation of this method includes the following steps: 1) Collect picture information in the scene through a visual sensor; 2) Input the picture information into a multi-modal large model to obtain scene semantic information; 3) Perceive the scene based on the scene semantic information; 4) Compare the current scene semantic information with preset interaction conditions to determine whether environmental interaction is required; 5) When the interaction conditions are met, perform corresponding interaction behaviors.

2. The scene perception and interaction method based on a multimodal large model according to claim 1, wherein The multi-modal large model includes a multi-modal feature extraction model with pictures as input and feature vectors as output; The multi-modal large model includes a multi-modal large language model with pictures and semantic information prompt words as input.

3. The scene perception and interaction method based on a multi-modal large model according to claim 2, wherein, The multi-modal feature extraction model with pictures as input and feature vectors as output includes a CLIP model; The multi-modal large language model with pictures and semantic information prompt words as input includes LLAVA and miniCPM methods; The semantic information prompt words are descriptions of what types of information can be obtained from pictures, including description prompt words for general information and description prompt words for specific information; the description prompt words for general information are prompt words for extracting information in all scenes; the description prompt words for specific information are prompt words formulated for certain special perception needs and interaction needs.

4. A scene perception and interaction method based on a multimodal large model according to claim 1 or 2 or 3, characterized in that, The scene semantic information includes the feature vectors extracted by the multi-modal feature extraction model and the image descriptions extracted by using the multi-modal large language model.

5. A scene perception and interaction method based on a multimodal large model according to claim 1, characterized in that The perceiving the scene based on the scene semantic information includes: Detecting and recording static objects in the environment, including real-time monitoring and recording of object categories, object positions, environmental states, and scene changes; Detecting and recording dynamic objects in the environment, including monitoring and recording of personnel categories, the number of personnel, personnel behaviors, and machine operating states.

6. The scene perception and interaction method based on a multimodal large model according to claim 1, wherein, The interaction conditions include abnormal interaction conditions and normal interaction conditions; The abnormal interaction conditions include static object abnormalities, including object abnormalities and environmental abnormalities; and also include dynamic object abnormalities, including personnel abnormalities and machine abnormalities; The normal interaction conditions include the reporting conditions for the current environmental semantic information, the personnel interaction conditions when there are people in the environment, and the conditions for reporting the object states in the environment; In step 4), When the abnormal interaction conditions are met, the interaction behaviors for the abnormal interaction conditions include reporting and reminding the abnormal state within the robot system, real-time reporting of the abnormality, voice reminder and intervention for personnel abnormalities, reporting and robotic arm operation intervention for machine abnormalities, and robotic arm operation intervention for abnormalities in the environmental state; The interaction behaviors for meeting the normal interaction conditions include reporting and reminding the current scene semantic information within the robot system, initiating an interaction invitation to the people in the environment through voice, and evaluating and reporting the status based on the objects in the environment.

7. A scene perception and interaction method based on a multimodal large model according to claim 1 or 6, characterized in that, In step 5), the interaction behaviors include grasping and moving objects in the environment based on the interaction content and navigation planning actions.

8. A scene perception and interaction system based on a multimodal large model, characterized in that, It includes: A visual sensor for collecting picture information in the scene; A multi-modal large model for processing the picture information and extracting scene semantic information; A perception module for perceiving the scene based on the scene semantic information; An interaction determination module for comparing the current scene semantic information with preset interaction conditions; An interaction execution module for performing corresponding interaction behaviors when the interaction conditions are met; The multi-modal large model supports real-time computing to ensure the rapid response of the robot in a dynamic environment; The perception module can dynamically record and retrieve objects, environments, and states in the scene; The interaction execution module supports various interaction behaviors, including voice broadcast, visual grasping, and navigation planning; This system can implement the scene perception and interaction method based on the multi-modal large model described in any one of claims 1 to 7.

9. A scene perception and interaction device based on a multimodal large model, characterized in that Including: At least one memory and at least one processor; The at least one memory is used to store machine-readable programs; The at least one processor is used to call the machine-readable program to implement the method described in any one of claims 1 to 7.

10. A computer-readable medium, characterized in that, Computer instructions are stored on the computer-readable medium, and when the computer instructions are executed by the processor, the method described in any one of claims 1 to 7 is implemented.