Method and device for enhancing spatial perception capability of multimodal large models
Image and video features are extracted through multimodal large model, grid processing and semantic map construction, and coordinate information of multimodal large model is optimized, which solves the problem of insufficient spatial positioning capabilities of multimodal large models in traditional methods, and realizes precise positioning in dynamic scenarios and target recognition in complex scenarios.
Patent Information
- Application Number
- CN202510792038.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-06-13
AI Technical Summary
Traditional vision methods and vision transformers limit the mapping and prediction capabilities of the precise coordinates of multimodal large models, resulting in insufficient spatial positioning capabilities of multimodal large models, increasing learning complexity, poor adaptability of dynamic scenes, and limiting their performance in actual precise spatial positioning tasks.
The multimodal big model extracts the feature description information of the object in the target image and/or video, generates initial structured data, and performs grid processing to add visual prompts for position information. Combining semantic maps and multi-frame memory mechanisms, the object coordinate information is optimized, and the actual object coordinates are finally mapped back to the system coordinates for precise positioning.
It improves the spatial understanding and precise positioning capabilities of multimodal large models in dynamic scenarios, reduces the complexity of intelligent scene positioning tasks, supports diverse interactions and output forms, is highly adaptable, and expands application scenarios.
Smart Images

Figure CN120339399B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision technology, and in particular to a method and device for enhancing the spatial perception capability of a multimodal large model. Background Art
[0002] Among related technologies, traditional vision models offer efficient processing speeds and low computational resource requirements for object detection and classification tasks, making them suitable for applications with high real-time requirements. They can accurately locate objects and quickly classify them in fixed environments, and are widely used in fields such as target detection and autonomous driving. However, these traditional vision models have weak generalization capabilities and typically require extensive training on specific datasets. Their performance often degrades significantly when faced with new scenes or unseen objects, making them difficult to cope with complex dynamic environments. Multimodal Large Models (MLMs) have made significant progress in recent years at the interface between computer vision and natural language processing. By combining visual and language modalities, these models enable complex semantic understanding and language interaction of visual content. Models like CLIP, through cross-modal embedding and alignment mechanisms, have demonstrated excellent zero-shot capabilities in multiple tasks.
[0003] However, in related technologies, traditional visual methods mostly rely on visual feature extraction backbone networks to obtain the semantic features of images. Although these methods work well in object recognition, the abstraction of spatial information in the deep feature extraction process leads to a significant decline in the precise coordinate mapping capability of large multimodal models. Secondly, the Vision Transformer (ViT) maintains spatial relationships by dividing the image into small blocks and using position encoding. The position encoding of this method is usually implicit, and large multimodal models need to learn to understand this encoded information, which makes it difficult to predict precise coordinates. This increases the complexity of learning large multimodal models and limits their performance in precise spatial positioning tasks, which needs to be urgently addressed. Summary of the Invention
[0004] The present application provides a method and device for enhancing the spatial perception capability of a large multimodal model to address the problems in related technologies, such as traditional visual methods and visual transformers limiting the mapping and prediction capabilities of the precise coordinates of a large multimodal model, resulting in deficiencies in the spatial positioning capability of the large multimodal model, and increasing the learning complexity of the large multimodal model, resulting in poor adaptability to dynamic scenes and limited application flexibility of the large multimodal model, which limits its performance in actual precise spatial positioning tasks.
[0005] The first aspect of the present application provides a method for enhancing the spatial perception capability of a multimodal large model, comprising the following steps: extracting feature description information of at least one object in a target image and / or a target video using a multimodal large model, and generating initial structured data of the at least one object based on the feature description information; performing gridding processing on the target image and / or the target video, and adding visual cues containing location information in multiple grids of the target image and / or the target video, so as to combine the visual cues and the initial structured data to generate structured data containing coordinate information and description information of the at least one object; locating a target area corresponding to the at least one object in the target image and / or the target video based on the structured data, and optimizing the coordinate information through the target area to obtain the actual object coordinates of the at least one object; mapping the actual object coordinates back to the system coordinates of the target image and / or the target video to obtain the actual positioning result of the at least one object in space.
[0006] Optionally, in one embodiment of the present application, the use of a multimodal large model to extract feature description information of at least one object in a target image and / or target video, and generating initial structured data of the at least one object based on the feature description information, includes: using the multimodal large model to convert the format of the initial target image and / or the initial target video to obtain the target image and / or target video, and obtaining standardized format data containing semantic information and position data of the at least one object according to user needs; based on the target image and / or target video, recording change information of the at least one object to generate a time series semantic description; based on the standardized format data, extracting the category, description, relative position information of the at least one object and its spatial association with surrounding objects to determine the initial structured data of the at least one object in combination with user needs and the time series semantic description.
[0007] Optionally, in one embodiment of the present application, the combining of the visual cues and the initial structured data to generate structured data containing coordinate information and description information of the at least one object includes: constructing a semantic map based on the initial structured data and the time series semantic description; and determining the structured data based on the semantic map and the visual cues.
[0008] Optionally, in one embodiment of the present application, the target image and / or the target video is gridded to add visual cues containing location information in multiple grids of the target image and / or the target video, including: determining specification information of the grid processing based on user needs, so as to grid the target image and / or the target video according to the specification information; and labeling the multiple grids according to mapping information between the multiple grids and the actual area to complete the adding of the visual cues.
[0009] Optionally, in one embodiment of the present application, optimizing the coordinate information through the target area to obtain the actual object coordinates of the at least one object includes: generating a refined grid on the image or video frame corresponding to the target area; determining the refined spatial coordinates of the at least one object in the refined grid to optimize the coordinate information and determine the actual object coordinates of the at least one object.
[0010] Optionally, in one embodiment of the present application, it also includes: judging whether the actual positioning result meets the user's needs; if the actual positioning result meets the user's needs, using the multimodal large model to transmit the structured data corresponding to the actual positioning result to the target user, otherwise, correcting the actual positioning result according to the user's needs to meet the user's needs.
[0011] The second aspect of the present application provides an apparatus for enhancing the spatial perception capability of a multimodal large model, comprising: an extraction module for extracting feature description information of at least one object in a target image and / or a target video using a multimodal large model, and generating initial structured data of the at least one object based on the feature description information; a processing module for performing gridding processing on the target image and / or the target video, and adding visual cues containing location information in multiple grids of the target image and / or the target video, so as to combine the visual cues and the initial structured data to generate structured data containing coordinate information and description information of the at least one object; a positioning module for locating a target area corresponding to the at least one object in the target image and / or the target video based on the structured data, and optimizing the coordinate information through the target area to obtain the actual object coordinates of the at least one object; a mapping module for mapping the actual object coordinates back to the system coordinates of the target image and / or the target video to obtain the actual positioning result of the at least one object in space.
[0012] Optionally, in one embodiment of the present application, the extraction module includes: a conversion unit for using the multimodal large model to perform format conversion on the initial target image and / or the initial target video to obtain the target image and / or target video, and obtain standardized format data containing semantic information and position data of the at least one object according to user needs; a first generation unit for recording change information of the at least one object based on the target image and / or target video to generate a time series semantic description; an extraction unit for extracting the category, description, relative position information of the at least one object and its spatial association with surrounding objects based on the standardized format data to determine the initial structured data of the at least one object in combination with user needs and the time series semantic description.
[0013] Optionally, in one embodiment of the present application, the processing module includes: a construction unit for constructing a semantic map based on the initial structured data and the time series semantic description; and a determination unit for determining the structured data based on the semantic map and the visual prompt.
[0014] Optionally, in one embodiment of the present application, the processing module includes: a first processing unit, used to determine the specification information of the grid processing based on user needs, so as to perform grid processing on the target image and / or the target video according to the specification information; a second processing unit, used to label the multiple grids according to the mapping information between the multiple grids and the actual area, so as to complete the addition of the visual prompts.
[0015] Optionally, in one embodiment of the present application, the positioning module includes: a second generation unit, used to generate a refined grid on the image or video frame corresponding to the target area; an optimization unit, used to determine the refined spatial coordinates of the at least one object in the refined grid to optimize the coordinate information and determine the actual object coordinates of the at least one object.
[0016] Optionally, in one embodiment of the present application, it also includes: a judgment module, used to judge whether the actual positioning result meets the user requirements; a correction module, used to use the multimodal large model to transmit the structured data corresponding to the actual positioning result to the target user when the actual positioning result meets the user requirements; otherwise, the actual positioning result is corrected according to the user requirements to meet the user requirements.
[0017] The third aspect of the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the program to implement the method for enhancing the spatial perception capability of a multimodal large model as described in the above embodiment.
[0018] The fourth aspect of the present application provides a computer-readable storage medium, which stores a computer program. When the program is executed by a processor, it implements the above method for enhancing the spatial perception capability of a multimodal large model.
[0019] The fifth aspect of the present application provides a computer program product, including a computer program, which, when executed, is used to implement the above method for enhancing the spatial perception capability of a multimodal large model.
[0020] The embodiment of the present application can use a multimodal large model to extract feature description information of objects in a target image and / or target video and generate initial structured data; then the target image and / or target video are gridded and visual cues containing location information are added to the grid to generate structured data containing coordinate information and description information of the object; then the actual object coordinates of at least one object are obtained by refining the target area corresponding to at least one object; finally, the actual object coordinates are mapped back to the system coordinates of the target image and / or target video to obtain the actual positioning result of at least one object in space. Thus, by combining multimodal semantic extraction, explicit gridding positioning, multi-frame memory mechanism and semantic map construction, the spatial understanding and precise positioning ability of the multimodal large model in dynamic scenes are improved, and target recognition and precise positioning from coarse-grained to fine-grained in complex scenes are completed; and the explicit gridding positioning and semantic map construction in the present application can greatly reduce the complexity of the multimodal large model in the intelligent scene positioning task. At the same time, the present application supports a variety of interaction and output forms, is highly adaptable, effectively expands the application scenarios of the multimodal large model, and can meet the needs of various application scenarios. This solves the problem in related technologies that traditional visual methods and visual transformers limit the mapping and prediction capabilities of the precise coordinates of the multimodal large model, resulting in insufficient spatial positioning capabilities of the multimodal large model, and increasing the learning complexity of the multimodal large model, making the multimodal large model poorly adaptable to dynamic scenes and limited in application flexibility, thereby limiting its performance in actual precise spatial positioning tasks.
[0021] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0023] Figure 1 A flowchart of a method for enhancing the spatial perception capability of a multimodal large model provided according to an embodiment of the present application;
[0024] Figure 2 This is a flowchart of dynamic updating and global analysis of a semantic map according to one embodiment of the present application;
[0025] Figure 3 This is a flow chart of a method for enhancing the spatial positioning capability of a multimodal large model according to one embodiment of the present application;
[0026] Figure 4 This is a block diagram of the modules for enhancing the spatial perception capability of a multimodal large model according to an embodiment of the present application;
[0027] Figure 5 This is a flowchart of an embodiment of the present application for performing image-based target recognition and positioning;
[0028] Figure 6 This is a schematic diagram of the execution flow of the first visual positioning function according to one embodiment of the present application;
[0029] Figure 7 This is a schematic diagram of the execution flow of the second visual positioning function according to one embodiment of the present application;
[0030] Figure 8 This is a schematic diagram of an example result in the field of robot operation according to one embodiment of the present application;
[0031] Figure 9 This is a flowchart of object positioning and operation based on user input for visual question answering according to one embodiment of the present application;
[0032] Figure 10 Schematic diagram of the structure of an apparatus for enhancing the spatial perception capability of a multimodal large model according to an embodiment of the present application;
[0033] Figure 11 A schematic diagram of the structure of an electronic device provided according to an embodiment of the present application.
[0034] Reference numerals:
[0035] 10-Device for enhancing the spatial perception capability of a multimodal large model: 100-extraction module, 200-processing module, 300-positioning module, 400-mapping module; 1101-memory, 1102-processor, 1103-communication interface. DETAILED DESCRIPTION
[0036] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0037] The following describes a method and device for enhancing the spatial perception capability of a multimodal large model according to an embodiment of the present application with reference to the accompanying drawings. In view of the related technologies mentioned in the above background technology, traditional visual methods and visual transformers limit the mapping and prediction capabilities of the precise coordinates of the multimodal large model, resulting in deficiencies in the spatial positioning capability of the multimodal large model, and increasing the learning complexity of the multimodal large model, making the multimodal large model poorly adaptable to dynamic scenes and having limited application flexibility, which limits its performance in actual precise spatial positioning tasks. The present application provides a method for enhancing the spatial perception capability of a multimodal large model, in which the multimodal large model can be used to extract feature description information of objects in a target image and / or target video and generate initial structured data; then the target image and / or target video are gridded and visual cues containing location information are added to the grid to generate structured data containing the coordinate information and description information of the object; then the actual object coordinates of the at least one object are obtained by refining the target area corresponding to the at least one object; finally, the actual object coordinates are mapped back to the system coordinates of the target image and / or target video to obtain the actual positioning result of the at least one object in space. Thus, it is achieved that by combining multimodal semantic extraction, explicit grid positioning, multi-frame memory mechanism and semantic map construction, the spatial understanding and precise positioning capabilities of the multimodal large model in dynamic scenes are improved, and target recognition and precise positioning from coarse-grained to fine-grained in complex scenes are completed; and the explicit grid positioning and semantic map construction in this application can greatly reduce the complexity of the multimodal large model in the intelligent scene positioning task. At the same time, this application supports a variety of interaction and output forms, is highly adaptable, and effectively expands the application scenarios of the multimodal large model to meet the needs of various application scenarios. Thus, it solves the problem in the related technology that traditional visual methods and visual transformers limit the mapping and prediction capabilities of the precise coordinates of the multimodal large model, resulting in deficiencies in the spatial positioning capability of the multimodal large model, and increases the learning complexity of the multimodal large model, making the multimodal large model poorly adaptable to dynamic scenes and limited in application flexibility, limiting its performance in actual precise spatial positioning tasks.
[0038] Specifically, Figure 1 A flowchart of a method for enhancing the spatial perception capability of a multimodal large model provided in an embodiment of the present application.
[0039] like Figure 1 As shown, the method for enhancing the spatial perception capability of a multimodal large model includes the following steps:
[0040] In step S101 , feature description information of at least one object in a target image and / or target video is extracted using a multimodal large model, and initial structured data of the at least one object is generated based on the feature description information.
[0041] It is understood that the target images and target videos here refer to images and videos received and processed by the multimodal large model when performing precise spatial positioning tasks in actual applications. These can be, but are not limited to, images and videos received through multimodal interaction with the multimodal large model via a user's interactive terminal, such as natural language conversation, voice input, image or video annotation, and gesture commands, and then processed by the multimodal large model itself.
[0042] In some embodiments, when enhancing the spatial perception capability of the multimodal large model, that is, enhancing the precise spatial positioning capability of the multimodal large model in practical applications, the present application can, but is not limited to, first use the multimodal large model to extract feature description information of at least one object in the target image or target video or the target image and target video, so as to further accurately locate the at least one object based on the feature description information.
[0043] It should be noted that, at least one object here can be understood as an object that the multimodal large model needs to accurately locate. It is at least one, but it can also be multiple. For example, the user inputs an image of a bookshelf and wants to find one, two, three, etc. books he needs from the image. The specific details can be determined by professional and technical personnel in this field or users based on actual application scenarios and actual needs. The embodiments of this application are only for illustrative purposes and are not specifically limited.
[0044] Also, the target image and / or target video in the subsequent process may be understood as the target image or the target video or the target image and the target video.
[0045] Furthermore, in order to reduce the processing complexity of the multimodal large model, the embodiment of the present application can also convert these feature description information into certain structured data to facilitate subsequent processing.
[0046] Optionally, in one embodiment of the present application, a multimodal large model is used to extract feature description information of at least one object in a target image and / or a target video, and initial structured data of at least one object is generated based on the feature description information, including: using the multimodal large model to convert the format of the initial target image and / or the initial target video to obtain the target image and / or the target video, and obtaining standardized format data containing semantic information and position data of at least one object according to user needs; based on the target image and / or the target video, recording change information of at least one object to generate a time series semantic description; based on the standardized format data, extracting the category, description, relative position information of at least one object and its spatial association with surrounding objects to determine the initial structured data of at least one object in combination with user needs and the time series semantic description.
[0047] It can be understood that the initial target image and / or initial target video here can be understood as an unprocessed image or video or an image and video input by the user.
[0048] During the actual execution process, when extracting feature description information of at least one object, in order to enable the image or video or image and video input by the user to be processed quickly and efficiently in the multimodal large model, the present application can first convert the format of the initial target image and / or the initial target video to obtain the target image and / or target video.
[0049] Specifically, the user can input the initial target image and / or initial target video into the multimodal large model. After receiving the target image and / or target video, the multimodal large model pre-processes the target image and / or target video, that is, uses image coding technology to convert the format of the target image and / or target video, including but not limited to image format conversion, noise removal, contrast adjustment and other steps to ensure that the image can adapt to subsequent processing, thereby providing the necessary basis for further analysis and positioning.
[0050] At the same time, considering the different scenario requirements of users, the multimodal large model has a high ability to process image information, but its text comprehension ability may be insufficient. Therefore, the embodiment of the present application can also introduce a text processing module to extract text from the image to determine the user's demand information, so as to obtain standardized format data containing semantic information and location data of at least one object according to user needs.
[0051] For example, a user might input an image of a bookshelf and want to find a specific book within it. Besides the text on the book cover, the image might also contain text such as "Please help me find XX book." The text processing module can extract this information. The standardized format of the data is: User Requirement: XXX; Target Object Name: XXX; Target Object Location: XXX, etc. This standardized format can serve as a reference for the output format of structured data in subsequent processes.
[0052] Therefore, the embodiment of the present application can realize the joint processing of user needs and image data using standardized processes, thereby generating standardized format data that meets standardization requirements and contains semantic information and location data of at least one object as input for subsequent steps.
[0053] Next, the embodiment of the present application can be based on the target image and / or target video, combined with a multi-frame memory mechanism, to record the object change information of at least one object in a dynamic scene, and generate a time series semantic description to adapt to the needs of the dynamic environment. That is, the embodiment of the present application can record the change information of at least one object in the scene in combination with multi-frame image data, thereby determining the position of the object, thereby effectively reducing the positioning error. For example, in 100 frames of image data, object A in 99 frames of images is on the desk, and in 1 frame of image, object A is on the chair, then it can be determined that the position of object A has changed, and is located in different positions at different times.
[0054] Finally, embodiments of the present application can utilize a multimodal large model based on standardized data formats to perform semantic analysis of images using a combination of visual and linguistic modalities. By combining the visual features of objects with linguistic modalities, the model can identify and extract feature description information for at least one object in the image, including but not limited to the object's category, description, spatial association with surrounding objects, and the relative positional relationship between the two. Furthermore, by combining user needs with time-series semantic descriptions, initial structured data containing semantic information and location data for at least one object can be generated.
[0055] Semantic information refers to the category and description within the feature description information. Position data refers to the initial location of an object, determined based on its spatial relationship with surrounding objects and their relative positional relationship, using an explicit spatial relative position reference. For example, object A is located on top of a desk. Initial structured data can be simply understood as structured object information description data, including but not limited to semantic information and position data for at least one object.
[0056] The embodiments of the present application can convert the original image information into a standardized format and generate initial structured data containing semantic information and position data of at least one object, providing an important basis for subsequent positioning and recognition tasks.
[0057] Step S102: gridding the target image and / or target video, and adding visual cues containing location information to multiple grids of the target image and / or target video, so as to combine the visual cues and the initial structured data to generate structured data containing coordinate information and description information of at least one object.
[0058] In some embodiments, the present application may perform gridding processing on the target image and / or target video so as to add visual cues containing location information in multiple grids of the target image and / or target video, and then use a multimodal large model to generate structured data containing object coordinate information and description information.
[0059] Among them, adding visual cues containing position information in multiple grids of the target image and / or target video can be understood as adding visual elements or marks that can represent the position coordinates of the grid in the image to the grids that have been divided into the target image and / or target video.
[0060] Then, by combining this visual cue with the initial structured data, which previously only contained semantic information and preliminary position data of the at least one object, structured data containing coordinate information and description information of the at least one object can be generated. The coordinate information here can be understood as the spatial coordinates of the at least one object in the target image and / or target video, obtained based on the visual cues containing position information in multiple grids of the target image and / or target video. The description information includes all the information in the initial structured data.
[0061] Furthermore, structured data containing coordinate information and description information for at least one object can be understood as structured data that adds relatively precise coordinate information of at least one object within a grid of a target image and / or target video, in addition to the original structured data containing semantic information and location data. This data combines the object's attributes and spatial relationships, helping to improve object recognition accuracy.
[0062] Next, the process of gridding the target image and / or target video in the embodiments of the present application to add visual cues containing location information in multiple grids of the target image and / or target video, and combining the visual cues and initial structured data to generate structured data containing coordinate information and description information of at least one object is further explained.
[0063] Optionally, in one embodiment of the present application, the target image and / or target video is gridded to add visual cues containing location information in multiple grids of the target image and / or target video, including: determining specification information of the grid processing based on user needs, so as to grid the target image and / or target video according to the specification information; and labeling the multiple grids according to mapping information between the multiple grids and the actual area to complete the addition of visual cues.
[0064] In certain embodiments, when performing gridding processing on a target image and / or target video, the target image and / or target video is divided into multiple grids of a fixed size. Considering that users have different requirements for accuracy and response speed, the present application can determine grid processing specifications based on user needs, thereby dynamically adjusting the grids based on this specifications.
[0065] Among them, the specification information of grid processing includes but is not limited to grid size, numbering rules or positioning accuracy, so as to ensure that the grid processing in the embodiment of the present application can adapt to complex scenarios and different task requirements.
[0066] Furthermore, embodiments of the present application can label multiple grids based on their mapping information to the actual area, i.e., label each grid with a unique number or coordinate. For example, if an image is divided into a 100*100 grid, the target area corresponding to the position of at least one object in the actual area is the grid at the intersection of the 66th column from left to right and the 66th row from bottom to top in the grid, and the label information can be 66′66, etc. This completes the addition of visual cues containing location information to multiple grids in the target image and / or target video.
[0067] It should be noted that professionals in this field can determine the type of visual cues based on actual application scenarios and actual needs, including but not limited to bounding boxes, markers, pixel-level cues, and soft cues. These visual cues can provide additional information to guide the model's interpretation and processing of visual data, thereby enhancing the model's visual perception capabilities.
[0068] The embodiments of the present application can generate visual cues containing location information within the grid, ensuring that each grid is associated with the location information of the area in which it is located. At the same time, it makes it easier for the multimodal large model to understand this encoded information, thereby improving the multimodal large model's ability to predict the precise coordinates of objects, and further improving the multimodal large model's performance in complex understanding and reasoning tasks. It can also provide an identification basis for precise positioning, reduce the learning complexity of the multimodal large model, and facilitate object positioning and subsequent precise labeling.
[0069] Optionally, in one embodiment of the present application, structured data containing coordinate information and description information of at least one object is generated in combination with visual cues and initial structured data, including: constructing a semantic map based on the initial structured data and time series semantic description; and determining structured data based on the semantic map and visual cues.
[0070] In other embodiments, after gridding the target image and / or target video and adding visual cues containing location information to the grid, the present application can input the gridded image into a multimodal large model. The multimodal large model can construct a semantic map based on the initial structured data and the time series semantic description, and then determine the structured data of at least one object based on the semantic map and the visual cues. In this way, structured data containing the coordinate information and description information of at least one object can be generated by combining the visual cues and the initial structured data.
[0071] For example, this application can combine multi-frame image data to record the dynamic changes of objects in a time series, including continuous changes in position, status, and description, and then build a semantic map based on structured data and time series information to achieve global semantic data integration of multi-frame images.
[0072] Among them, the semantic map here refers to the dynamic structured data constructed by combining multi-frame image data, the spatial position information of the object and the time series semantic description, which is a form of expression of structured data in the embodiment of the present application. It can record and reflect the position changes, state changes and spatial relationships of objects with surrounding objects in real time. For example, the semantic map records the preliminary positioning results of the object in the form of a JSON file, including the name of the object, horizontal and vertical coordinate information, timestamp and spatial relationship with surrounding objects.
[0073] This semantic map not only contains the semantic information and location information of objects, but also stores the spatial associations between objects and other items, supports the maintenance of continuous time series data, and adapts to complex and changing environments. Through this semantic map, the embodiments of the present application can maintain the association between historical information and current frame data, ensuring the global understanding ability of the multimodal large model in dynamic scenes and ensuring the precise positioning ability of the multimodal large model in dynamic scenes.
[0074] Inputting the semantic map and the current frame image data into the multimodal large model together makes it easier for the multimodal large model to perform joint analysis, thereby generating more accurate object location coordinate information and semantic data.
[0075] It should be noted that the semantic map in the embodiment of the present application is dynamic, not fixed. For example, Figure 2 This is a flowchart of the dynamic update and global analysis of the semantic map of an embodiment of this application. Figure 2 As shown, after receiving the target video, the embodiment of the present application will process multiple frames in the video and use the multi-frame memory mechanism to integrate the information of the previous and next frames to generate a time series semantic map. Alternatively, after receiving multiple target images with a time series, the embodiment of the present application can also use the multi-frame memory mechanism to integrate the information of the previous and next frames to generate a time series semantic map.
[0076] After generating a time-series semantic map, this embodiment of the present application determines whether the scene has changed as multiple frames of imagery are added. If the scene has not changed, the semantic map information remains unchanged and the semantic map and the current frame of imagery are input into a large multimodal model for multimodal analysis, generating structured data for at least one object. This data includes, but is not limited to, the object's description, coordinates, and spatial relationship information for subsequent processing.
[0077] If the scene changes, the objects appearing in the current frame and their dynamic positions are added according to the scene change information, or the positions or states of objects in historical frames are modified, or objects that no longer exist in the semantic map are deleted, thereby ensuring that the semantic map can reflect the changes of objects in the dynamic environment in real time.
[0078] It should also be noted that when generating the initial structured data, the multimodal large model in the embodiment of the present application can also construct a semantic map in combination with multi-frame image data in order to preliminarily determine the position of at least one object. The initial structured data constructs a semantic map in order to preliminarily determine the position of the object, and the structured data constructs a semantic map in order to further determine the position of the object more accurately. The modification time of the semantic map is not specifically limited in the embodiment of the present application, and can be achieved in the process of generating the initial structured data and the structured data. It can be set or adjusted by professionals in this field according to actual conditions and actual needs.
[0079] Step S103 : Based on the structured data, locate a target area corresponding to at least one object in the target image and / or target video, and optimize coordinate information through the target area to obtain actual object coordinates of the at least one object.
[0080] As a possible implementation method, after gridding the target image and / or target video and adding visual cues containing location information to the grid to obtain structured data containing coordinate information and description information of at least one object, in order to further improve the accuracy of the spatial coordinates of at least one object, the embodiment of the present application can also re-locate the target area corresponding to at least one object in the target image and / or target video based on the structured data, and then further refine the spatial coordinates of at least one object within the target area, thereby optimizing the previous spatial coordinates and obtaining the actual object coordinates of at least one object.
[0081] The target region here refers to the area corresponding to the grid position of at least one object in the target image and / or target video after the target image and / or target video is divided into grids. For example, if an image is divided into a 100x100 grid, the target region corresponding to at least one object is the grid at the intersection of the 66th column from the left to the right and the 66th row from the bottom to the top, with the label information 66′66.
[0082] It should be noted that before further processing the target area, the embodiment of the present application will first determine whether the target area meets the positioning requirements, that is, whether it is determined that there is at least one object that needs to be located in the target area, or whether the target area meets the positioning requirements of at least one object, such as maintaining the integrity of the object in the target area. When the target area meets the positioning requirements, the system will continue with subsequent processing. If the target area does not meet the recognition requirements, the embodiment of the present application can adjust the target area based on the preliminary detection results and re-perform regional analysis to ensure the accuracy and efficiency of the final recognition.
[0083] Optionally, in one embodiment of the present application, the coordinate information of the target area is optimized to obtain the actual object coordinates of at least one object, including: generating a refined grid on the image or video frame corresponding to the target area; determining the refined spatial coordinates of at least one object in the refined grid to optimize the coordinate information and determine the actual object coordinates of at least one object.
[0084] In some embodiments, when optimizing the coordinate information of at least one object through a target area to obtain the actual object coordinates of at least one object, the present application can, but is not limited to, accurately locate the target area corresponding to at least one object in the target image and / or target video based on the name and coordinates of the object in the structured data, and then crop, enlarge and refine the target area to generate a refined grid, and then determine the refined spatial coordinates of at least one object in the refined grid to optimize the previous coordinate information, that is, use this refined spatial coordinate as the latest spatial coordinate, thereby determining the actual object coordinates of at least one object.
[0085] For example, the present application can crop the target area in the target image and / or target video based on the previous spatial coordinate information, that is, take the area corresponding to the grid where at least one object is located as an image as a whole, crop the image as a whole from the target image and / or target video and generate a refined grid on the image as a whole.
[0086] Then, the embodiment of the present application can call a multimodal large model to combine visual and language modalities to further extract the precise location, boundary information and spatial association of the target object with surrounding objects; combine semantic maps and time series information to analyze the dynamic changes of objects, optimize the positioning accuracy of the target area, and determine the refined spatial coordinates of at least one object.
[0087] Ultimately, the embodiment of the present application can generate accurate structured data containing object categories, descriptions, positions and dynamic relationships based on the refined spatial coordinates, thereby meeting high-precision positioning requirements in complex scenarios, especially dynamic scenarios.
[0088] It should be noted that the process of locating the target area corresponding to at least one object in the embodiment of the present application and then further refining the spatial coordinates of at least one object within the target area to optimize the previous spatial coordinates is not a one-time process, but can be iteratively optimized according to the actual needs and actual conditions of professional and technical personnel and users in this field until the accuracy requirements of the spatial positioning task or the user requirements are met.
[0089] Step S104 : Mapping the actual object coordinates back to the system coordinates of the target image and / or target video to obtain an actual positioning result of at least one object in space.
[0090] During the actual execution process, after obtaining the actual object coordinates of at least one object, the present application can map the actual object coordinates of the local image back to the system coordinates of the original target image and / or the original target video to obtain the actual positioning results of at least one object in the original target image and / or target video space.
[0091] In this process, to facilitate the output of the multimodal large model and user understanding, the embodiment of the present application can output the actual positioning results to the user in the form of structured data. The actual positioning results include but are not limited to the description, coordinates and spatial relationship information of the object.
[0092] Specifically, the embodiment of the present application can map the precise position of the cropped area back to the global coordinate system, and generate semantic data with dynamically changing labels in combination with time series; and output the semantic information of the object through the natural language generation module to clearly describe the category, location and status of the target object; finally, generate operation instructions or non-text tokens (such as action tokens) corresponding to user needs and output them to the user.
[0093] Action tokens are a specific type of non-text token used in robotics to represent specific actions or tasks that a robot should perform. By combining them with semantic map data, they can output results in the form of natural language descriptions, operational instructions, or precise coordinates, thereby improving the positioning accuracy and output flexibility of objects in complex scenes.
[0094] Optionally, in one embodiment of the present application, it also includes: judging whether the actual positioning result meets the user's needs; if the actual positioning result meets the user's needs, using the multimodal large model to transmit the structured data corresponding to the actual positioning result to the target user, otherwise the actual positioning result is corrected according to the user's needs to meet the user's needs.
[0095] In certain embodiments, after generating the actual positioning results, the present application will perform self-evaluation and detection on the generated results, such as verifying them in combination with confidence thresholds, historical frame records, etc. If a large deviation from expectations is found, secondary positioning or multi-frame correction will be triggered.
[0096] Furthermore, the embodiment of the present application will also determine whether the actual positioning results outputted meet user requirements. If so, the structured data corresponding to the actual positioning results will be transmitted to the target user using a multimodal large model, and the processing will be terminated. If not, adjustments will be made based on the analysis results to ensure that the final output meets user requirements.
[0097] For example, a user inputs a 20-second target video and is asked to find the time segment where a book appears. After receiving the localization task and the target video, the multimodal large model detects that a corner of the book appears in the upper right corner of the first 0.5 seconds of the video, the book does not appear between 0.5 and 3 seconds, and the entire book appears between 3 and 20 seconds. Therefore, the output localization result is for the video between 3 and 20 seconds, with the book located at the intersection of the 66th column from the left and the 66th row from the bottom of the 100x100 grid. However, the user requires localization results for all occurrences of the book, not just the complete book. Therefore, adjustments must be made based on the analysis results to ensure the final output meets the user's requirements.
[0098] Among them, the structured data output by the embodiment of the present application can be expressed in various forms, such as natural language description, precise coordinates, operation instructions or behavioral instructions based on special marks, etc., providing users with rich interaction methods and adapting to diverse application needs in complex scenarios.
[0099] The present application is explained in detail below using a specific embodiment.
[0100] Figure 3 This is a flow chart of a method for enhancing the spatial positioning capability of a multimodal large model according to one embodiment of the present application; Figure 4 This is a schematic diagram of the principle of a method for enhancing the spatial perception capability of a multimodal large model according to an embodiment of the present application. In this case, each module in the system can also be independently embedded in a multimodal large model. This is only an exemplary explanation without specific limitation. Figure 3 and Figure 4 As shown:
[0101] Step S301: Using the image information extraction function, the input image or video is preliminarily processed to extract the feature description information of the object and generate initial structured data, including the object category, description, and relative position relationship of the object;
[0102] Step S302 , using the first visual positioning function, gridding the input image or video, generating a preliminary recognition result including object location and description information, and converting the result into structured data including object spatial coordinates (coordinate information) and description information;
[0103] Step S303: Using the second visual positioning function, combined with the name and coordinates of the object (based on the object's location identified by the first visual positioning function), the target area is accurately located in the image, and the target area is cropped, enlarged, and meshed. The system is then used to further optimize the precise spatial coordinates of the object.
[0104] In step S304, the fine coordinates of the local image are mapped back to the system coordinates of the original image to generate structured data containing the precise positioning results of the object. Based on user needs, the result output function is used to output the data in various forms (such as natural language descriptions, precise coordinates, operation instructions, or behavioral instructions based on special tags) to facilitate subsequent analysis and application, while adapting to diverse application needs in complex scenarios.
[0105] Next, the embodiment of the present application is combined with Figure 5 、 Figure 6 and Figure 7 The specific execution process of steps S201 to S203 is further explained. Figure 5 This is a flowchart of an embodiment of the present application for performing image-based target recognition and positioning. Figure 6 This is a schematic diagram of the execution flow of the first visual positioning function of an embodiment of the present application. Figure 7 This is a schematic diagram of the execution flow of the second visual positioning function of an embodiment of the present application, such as Figure 5 、 Figure 6 and Figure 7 As shown:
[0106] First, the image is preprocessed to generate standardized formatted data containing semantic information and location data. Then, a multimodal large model is used to perform semantic analysis and extract the feature description information of the object to generate initial structured data, that is, structured object information description data.
[0107] Then, if Figure 6 As shown, the image is gridded through the first visual positioning function, and a unique number or coordinate information is added to each grid to help locate the relative position of the object; on this basis, according to the semantic information of the object, the multimodal large model is called to identify the object, and the position of the object is coarse-grained spatially positioned, generating a coarse-grained recognition result of structured output data containing the object position and description information, further supporting fine-grained image positioning.
[0108] Then, if Figure 7As shown in the figure, the second visual positioning function further processes the coarse-grained recognition results of the image to achieve precise object positioning. First, the coarse-grained image recognition results are used to extract the description and location information of a single object. Based on this information, the target object is segmented into sub-images in the original image for further processing. On this basis, the image is divided into multiple grids through gridding processing, and coding or coordinate information is added to each grid to assist in subsequent precise positioning. Next, the multimodal large model is called to recognize the gridded image and generate fine-grained object positioning results. This process ensures that the precise position of the target object can be accurately identified and provides accurate data support for subsequent tasks.
[0109] as well as, Figure 8 This is a schematic diagram of an example result in the field of robot operation according to one embodiment of the present application. Figure 8 As shown, the embodiment of the present application can achieve dynamic scene target recognition and precise capture through a multimodal large model. Among them, the multimodal large model can use, for example, GPT-4o or LLaMA 3.2 Vision, etc., and can be selected or adjusted by professionals in this field according to actual conditions. This is only an example and is not a specific limitation.
[0110] Specifically, users can express requests to the system using natural language commands, such as "I'm a little thirsty." The multimodal large model receives the command and performs semantic analysis to extract the specific action required, such as "grab the water cup." Based on this semantic analysis, the system uses visual sensors to capture an image of the current scene and performs preliminary image processing to extract basic feature descriptions of the objects, generating structured data including the object's category, description, and relative position.
[0111] The multimodal large model then performs coarse-grained recognition of target objects in the scene. By gridding the image, unique coordinate information is annotated for each grid, and based on the object's visual features and semantic information, the target object's category, description, and approximate location in the scene are identified. A semantic map is generated based on the recognition results.
[0112] Based on coarse-grained recognition, the system further crops the area where the target object is located and refines the grid of the cropped area to improve the positioning accuracy of the target object. At this time, the semantic map provides contextual information and global spatial association information to assist the system in more efficiently and accurately analyzing the cropped area. The multimodal large model is used to further process the cropped image to identify the boundary features, shape of the target object and its spatial relationship with the surrounding environment. Combining the semantic map and the detailed analysis results of the cropped area, the system is able to generate the precise position coordinates of the target object and map these coordinates to the global coordinate system of the original image. The generated output data includes the precise coordinates of the target object, a dynamic semantic description, and a spatial relationship with other objects, such as "the water cup is in the center of the table, close to the left side of the book."
[0113] Semantic maps not only play a role in the initial recognition phase but also dynamically update the time series information of objects, such as recording the object's movement trajectory or position changes. This ensures that even if the scene changes, the system can accurately capture the state of the target object and maintain real-time global semantic understanding.
[0114] Based on the precise location of the target object and the user's needs, the system generates corresponding operational instructions and sends them to the robot's execution module. For example, the instructions include the target object's coordinates and a specific action description, such as "grab the cup at position (x, y, z)." After receiving the instructions, the robot uses its robotic arm to perform positioning and grasping operations, delivering the cup to the user or the designated location. After the operation is completed, the system updates the semantic map and generates a task execution report, including task execution status, time, and final location confirmation, for further analysis by the user or the system.
[0115] As a result, the advantages of multimodal large models in visual semantic understanding and spatial positioning are combined. Through the integration of natural language interaction and robot dynamic operations, accurate object recognition and efficient grasping operations in complex scenes are achieved, which significantly improves the robot's scene adaptability in dynamic environments.
[0116] It should be noted that the method for enhancing the spatial perception capability of a large multimodal model in this application is not limited to the field of robotics. It can also be widely used in fields such as automatic control of smart homes, environmental perception and path planning in autonomous driving, target item management and grasping in intelligent warehousing systems, and precise detection of lesion areas in medical images. The specific implementation methods in different fields can be adjusted by professional and technical personnel in this field according to actual application scenarios and actual needs to achieve a comprehensive improvement in the spatial positioning capability of large multimodal models in this application while meeting diverse application needs.
[0117] Furthermore, the embodiments of the present application will make certain corrections to the output results. Figure 9 This is a flowchart of object positioning and operation based on user input for visual question answering (VQA) in one embodiment of the present application. Figure 9 As shown, after obtaining the object description information from the text description provided by the user, the embodiment of the present application will determine whether the object information exists in the semantic information. If the object information exists in the semantic information, the object positioning is continued based on the information; if the object information does not exist, the first visual positioning function is called to perform coarse-grained object recognition and extract the approximate position of the object. After coarse-grained recognition, the second visual positioning function is called to perform fine-grained object recognition to further accurately locate the position of the target object. The fine-grained recognition module extracts the precise position of the object through high-resolution image processing and precise grid division technology. Finally, based on the recognition result, the position of the object or the operation instruction is output. The output result can be used to control the device to perform precise operations on the target object, such as positioning, grasping, etc.
[0118] According to the method for enhancing the spatial perception capability of a multimodal large model proposed in an embodiment of the present application, the multimodal large model can be used to extract feature description information of objects in a target image and / or target video and generate initial structured data; the target image and / or target video is then gridded and visual cues containing location information are added to the grid, and structured data containing coordinate information and description information of the object is generated in combination with the visual cues; the actual object coordinates of at least one object are then obtained by refining the target area corresponding to at least one object; finally, the actual object coordinates are mapped back to the system coordinates of the target image and / or target video to obtain the actual positioning result of at least one object in space. Thus, it is achieved that by combining multimodal semantic extraction, explicit grid positioning, multi-frame memory mechanism and semantic map construction, the spatial understanding and precise positioning capabilities of the multimodal large model in dynamic scenes are improved, and target recognition and precise positioning from coarse-grained to fine-grained in complex scenes are completed; and the explicit grid positioning and semantic map construction in this application can greatly reduce the complexity of the multimodal large model in the intelligent scene positioning task. At the same time, this application supports a variety of interaction and output forms, is highly adaptable, and effectively expands the application scenarios of the multimodal large model to meet the needs of various application scenarios. Thus, it solves the problem in the related technology that traditional visual methods and visual transformers limit the mapping and prediction capabilities of the precise coordinates of the multimodal large model, resulting in deficiencies in the spatial positioning capability of the multimodal large model, and increases the learning complexity of the multimodal large model, making the multimodal large model poorly adaptable to dynamic scenes and limited in application flexibility, limiting its performance in actual precise spatial positioning tasks.
[0119] Next, a device for enhancing the spatial perception capability of a multimodal large model proposed in an embodiment of the present application will be described with reference to the accompanying drawings.
[0120] Figure 10 It is a structural diagram of the device for enhancing the spatial perception capability of a multimodal large model according to an embodiment of the present application.
[0121] like Figure 10 As shown, the device 10 for enhancing the spatial perception capability of a multimodal large model includes: an extraction module 100, a processing module 200, a positioning module 300, and a mapping module 400.
[0122] Among them, the extraction module 100 is used to extract feature description information of at least one object in the target image and / or target video using the multimodal large model, and generate initial structured data of the at least one object based on the feature description information.
[0123] The processing module 200 is used to perform gridding processing on the target image and / or target video, and add visual cues containing location information in multiple grids of the target image and / or target video, so as to combine the visual cues and the initial structured data to generate structured data containing coordinate information and description information of at least one object.
[0124] The positioning module 300 is used to locate a target area corresponding to at least one object in a target image and / or target video based on structured data, and optimize coordinate information through the target area to obtain actual object coordinates of the at least one object.
[0125] The mapping module 400 is configured to map the actual object coordinates back to the system coordinates of the target image and / or target video to obtain an actual positioning result of at least one object in space.
[0126] Optionally, in one embodiment of the present application, the extraction module 100 includes: a conversion unit, a first generation unit and an extraction unit.
[0127] The conversion unit is used to convert the target image and / or target video into a format so as to obtain standardized format data containing semantic information and position data of at least one object in accordance with user requirements.
[0128] The first generating unit is configured to record change information of at least one object based on a target image and / or a target video to generate a time series semantic description.
[0129] An extraction unit is used to extract the category, description, relative position information of at least one object and its spatial association with surrounding objects based on standardized format data, so as to determine the initial structured data of at least one object in combination with user needs and time series semantic description.
[0130] Optionally, in one embodiment of the present application, the processing module 200 includes: a construction unit and a determination unit.
[0131] Among them, the construction unit is used to build a semantic map based on the initial structured data and time series semantic description.
[0132] A determination unit, which is used to determine structured data based on semantic maps and visual cues.
[0133] Optionally, in one embodiment of the present application, the processing module 200 includes: a first processing unit and a second processing unit.
[0134] The first processing unit is configured to determine specification information of grid processing based on user requirements, so as to perform grid processing on the target image and / or target video according to the specification information.
[0135] The second processing unit is configured to perform labeling on the plurality of grids according to mapping information between the plurality of grids and the actual area, so as to complete the adding of visual cues.
[0136] Optionally, in one embodiment of the present application, the positioning module 300 includes: a second generating unit and an optimizing unit.
[0137] The second generating unit is configured to generate a refined grid on the image or video frame corresponding to the target area.
[0138] The optimization unit is configured to determine the refined spatial coordinates of at least one object in the refined grid to optimize the coordinate information and determine the actual object coordinates of the at least one object.
[0139] Optionally, in one embodiment of the present application, it further includes: a judgment module and a correction module.
[0140] The judgment module is used to judge whether the actual positioning result meets the user's needs.
[0141] The correction module is used to transmit the structured data corresponding to the actual positioning result to the target user using the multimodal large model when the actual positioning result meets the user's needs. Otherwise, the actual positioning result is corrected according to the user's needs to meet the user's needs.
[0142] It should be noted that the above explanation of the method embodiment for enhancing the spatial perception capability of a multimodal large model is also applicable to the device for enhancing the spatial perception capability of a multimodal large model in this embodiment, and will not be repeated here.
[0143] According to the device for enhancing the spatial perception capability of a multimodal large model proposed in an embodiment of the present application, the multimodal large model can be used to extract feature description information of objects in a target image and / or target video and generate initial structured data; the target image and / or target video is then gridded and visual cues containing location information are added to the grid to generate structured data containing coordinate information and description information of the object; the actual object coordinates of at least one object are then obtained by refining the target area corresponding to at least one object; finally, the actual object coordinates are mapped back to the system coordinates of the target image and / or target video to obtain the actual positioning result of at least one object in space. Thus, it is achieved that by combining multimodal semantic extraction, explicit grid positioning, multi-frame memory mechanism and semantic map construction, the spatial understanding and precise positioning capabilities of the multimodal large model in dynamic scenes are improved, and target recognition and precise positioning from coarse-grained to fine-grained in complex scenes are completed; and the explicit grid positioning and semantic map construction in this application can greatly reduce the complexity of the multimodal large model in the intelligent scene positioning task. At the same time, this application supports a variety of interaction and output forms, is highly adaptable, and effectively expands the application scenarios of the multimodal large model to meet the needs of various application scenarios. Thus, it solves the problem in the related technology that traditional visual methods and visual transformers limit the mapping and prediction capabilities of the precise coordinates of the multimodal large model, resulting in deficiencies in the spatial positioning capability of the multimodal large model, and increases the learning complexity of the multimodal large model, making the multimodal large model poorly adaptable to dynamic scenes and limited in application flexibility, limiting its performance in actual precise spatial positioning tasks.
[0144] Figure 11 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device may include:
[0145] A memory 1101 , a processor 1102 , and a computer program stored in the memory 1101 and executable on the processor 1102 .
[0146] When the processor 1102 executes the program, the method for enhancing the spatial perception capability of the multimodal large model provided in the above embodiment is implemented.
[0147] Furthermore, the electronic device further includes:
[0148] The communication interface 1103 is used for communication between the memory 1101 and the processor 1102 .
[0149] The memory 1101 is used to store computer programs that can be run on the processor 1102 .
[0150] The memory 1101 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0151] If the memory 1101, processor 1102, and communication interface 1103 are implemented independently, the communication interface 1103, memory 1101, and processor 1102 can be interconnected via a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be divided into address buses, data buses, control buses, etc. For ease of representation, Figure 11 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0152] Optionally, in a specific implementation, if the memory 1101, the processor 1102 and the communication interface 1103 are integrated on a chip, the memory 1101, the processor 1102 and the communication interface 1103 can communicate with each other through an internal interface.
[0153] The processor 1102 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0154] An embodiment of the present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above method for enhancing the spatial perception capability of a multimodal large model.
[0155] An embodiment of the present application also provides a computer program product, including a computer program, which can run computer instructions. When the computer instructions are executed by a processor, the method for enhancing the spatial perception capability of a multimodal large model provided in an embodiment of the present application is implemented.
[0156] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.
[0157] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Thus, a feature specified as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of this application, "N" means at least two, for example, two, three, etc., unless otherwise specifically defined.
[0158] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or N executable instructions for implementing a custom logical function or process step, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may be performed in a different order than shown or discussed, including performing functions in a substantially simultaneous manner or in a reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application pertain.
[0159] The logic and / or steps represented in a flowchart or otherwise described herein, for example, can be considered a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" is any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (not exhaustive) of computer-readable media include: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program can be obtained electronically by optically scanning the paper or other medium and then editing, interpreting or processing it in other suitable ways as necessary, and then storing it in a computer memory.
[0160] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiment, the N steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented using hardware, as in another embodiment, it can be implemented using any one or a combination of the following technologies known in the art: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having suitable combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0161] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0162] In addition, the functional units in the various embodiments of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into a module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0163] The storage medium mentioned above may be a read-only memory, a magnetic disk, or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present application. Persons skilled in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A method for enhancing the spatial perception capability of a multimodal large model, characterized in that: The following steps are involved: Extracting feature description information of at least one object in a target image and / or target video using the multimodal large model, and generating initial structured data of the at least one object based on the feature description information; performing gridding processing on the target image and / or the target video, and adding visual cues containing position information to multiple grids of the target image and / or the target video, so as to combine the visual cues with the initial structured data to generate structured data containing coordinate information and description information of the at least one object; Based on the structured data, locate a target area corresponding to the at least one object in the target image and / or the target video, and optimize the coordinate information using the target area to obtain actual object coordinates of the at least one object; The actual object coordinates are mapped back to the system coordinates of the target image and / or the target video to obtain an actual positioning result of the at least one object in space.
2. The method for enhancing the spatial perception capability of a multimodal large model according to claim 1, characterized in that: The extracting feature description information of at least one object in the target image and / or target video using the multimodal large model, and generating initial structured data of the at least one object based on the feature description information, includes: Performing format conversion on an initial target image and / or an initial target video using the multimodal large model to obtain the target image and / or target video, and obtaining standardized format data including semantic information and position data of the at least one object according to user requirements; Based on the target image and / or target video, record change information of the at least one object to generate a time series semantic description; Based on the standardized format data, the category, description, relative position information of the at least one object and its spatial association with surrounding objects are extracted to determine the initial structured data of the at least one object in combination with user needs and the time series semantic description.
3. The method for enhancing the spatial perception capability of a multimodal large model according to claim 2, characterized in that: The step of combining the visual prompt and the initial structured data to generate structured data including coordinate information and description information of the at least one object includes: Building a semantic map based on the initial structured data and the time series semantic description; The structured data is determined based on the semantic map and the visual cue.
4. The method for enhancing the spatial perception capability of a multimodal large model according to claim 1, characterized in that: The gridding process of the target image and / or the target video to add visual cues containing position information in multiple grids of the target image and / or the target video includes: Determining specification information of the grid processing based on user requirements, so as to perform grid processing on the target image and / or the target video according to the specification information; The multiple grids are labeled according to mapping information between the multiple grids and the actual area to complete the adding of the visual prompt.
5. The method for enhancing the spatial perception capability of a multimodal large model according to claim 1, characterized in that: Optimizing the coordinate information by using the target area to obtain actual object coordinates of the at least one object includes: Generating a refined grid on the image or video frame corresponding to the target area; Refined spatial coordinates of the at least one object are determined in the refined grid to optimize the coordinate information and determine actual object coordinates of the at least one object.
6. The method for enhancing the spatial perception capability of a multimodal large model according to claim 1, characterized in that: Also includes: Determining whether the actual positioning result meets user needs; If the actual positioning result meets the user's needs, the structured data corresponding to the actual positioning result is transmitted to the target user using the multimodal large model; otherwise, the actual positioning result is modified according to the user's needs to meet the user's needs.
7. A device for enhancing the spatial perception capability of a multimodal large model, characterized in that: include: an extraction module, configured to extract feature description information of at least one object in a target image and / or target video using the multimodal large model, and generate initial structured data of the at least one object based on the feature description information; a processing module, configured to perform gridding processing on the target image and / or the target video, and add visual cues containing position information to multiple grids of the target image and / or the target video, so as to combine the visual cues with the initial structured data to generate structured data containing coordinate information and description information of the at least one object; a positioning module, configured to locate a target area corresponding to the at least one object in the target image and / or the target video based on the structured data, and optimize the coordinate information using the target area to obtain actual object coordinates of the at least one object; A mapping module is used to map the actual object coordinates back to the system coordinates of the target image and / or the target video to obtain the actual positioning result of the at least one object in space.
8. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method for enhancing the spatial perception capability of a multimodal large model as described in any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the method for enhancing the spatial perception capability of a multimodal large model as described in any one of claims 1 to 6.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed, it is used to implement the method for enhancing the spatial perception capability of a multimodal large model as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Defect labeling method based on computer vision large model
CN116912830A
Method, model and device for enhancing visual perception capability of multi-modal large language model
CN118585954A