Method and device for enhancing spatial perception capability of multi-modal large model
Through the methods of multimodal large model extraction, grid processing and visual cues, the problem of insufficient spatial positioning capabilities of multimodal large models is solved, precise positioning in dynamic scenarios and target recognition in complex scenarios is achieved, and the adaptability and application flexibility of multimodal large models are improved.
Patent Information
- Application Number
- CN202510792038.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-13
AI Technical Summary
Traditional vision methods and vision transformers limit the mapping and prediction capabilities of the precise coordinates of multimodal large models, resulting in insufficient spatial positioning capabilities of multimodal large models, increasing learning complexity, poor adaptability of dynamic scenes, and limited application flexibility, which limits its performance in actual precise spatial positioning tasks.
The multimodal big model extracts the feature description information of the object in the target image and/or video, generates initial structured data, and performs grid processing, adds visual prompts of position information, combines visual prompts and initial structured data to generate structured data containing the coordinate information of the object and description information, optimizes the coordinate information through the target area, and finally maps the actual object coordinates back to the system coordinates to achieve precise positioning.
It improves the spatial understanding and precise positioning capabilities of multimodal large models in dynamic scenarios, reduces the complexity of intelligent scene positioning tasks, supports diverse interactions and output forms, is highly adaptable, expands application scenarios, and meets multiple needs.
Smart Images

Figure CN120339399A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision technology, and in particular to a method and device for enhancing the spatial perception capability of a large multimodal model. Background Art
[0002] In related technologies, traditional visual models have efficient processing speed and low computing resource requirements in object detection and classification tasks, and are suitable for application scenarios with high real-time requirements. It can accurately locate objects and quickly classify them in a fixed environment, and is widely used in fields such as target detection and autonomous driving. However, these traditional visual models have weak generalization capabilities and usually require a lot of training on specific data sets. When faced with new scenes or unseen objects, their performance often drops significantly, making it difficult to cope with complex dynamic environments. Multimodal Large Models (MLMs) have made significant progress in the field of interaction between computer vision and natural language processing in recent years. By combining visual and language modalities, these models can perform complex semantic understanding and language interaction on visual content. Models such as CLIP have demonstrated excellent zero-shot capabilities in multiple tasks through cross-modal embedding and alignment mechanisms.
[0003] However, in the related technologies, traditional visual methods mostly rely on visual feature extraction backbone networks to obtain the semantic features of images. Although these methods work well in object recognition, the abstraction of spatial information in the deep feature extraction process leads to a significant decrease in the precise coordinate mapping capability of large multimodal models. Secondly, the Vision Transformer (ViT) divides the image into small blocks and uses position encoding to maintain spatial relationships. The position encoding of this method is usually implicit, and large multimodal models need to learn to understand this encoded information, which makes it difficult to predict precise coordinates. This increases the complexity of learning large multimodal models and limits their performance in precise spatial positioning tasks, which needs to be solved urgently. Summary of the invention
[0004] The present application provides a method and device for enhancing the spatial perception capability of a large multimodal model, so as to solve the problem in the related art that traditional visual methods and visual transformers limit the mapping and prediction capabilities of the precise coordinates of the large multimodal model, resulting in deficiencies in the spatial positioning capability of the large multimodal model, and increasing the learning complexity of the large multimodal model, resulting in poor adaptability to dynamic scenes and limited application flexibility of the large multimodal model, which limits its performance in actual precise spatial positioning tasks.
[0005] An embodiment of the first aspect of the present application provides a method for enhancing the spatial perception ability of a multimodal large model, including the following steps: extracting feature description information of at least one object in a target image and / or a target video by using the multimodal large model, and generating initial structured data of the at least one object according to the feature description information; performing grid processing on the target image and / or the target video, and adding visual cues containing position information to multiple grids of the target image and / or the target video, so as to generate structured data containing coordinate information and description information of the at least one object by combining the visual cues and the initial structured data; based on the structured data, locating a target area corresponding to the at least one object in the target image and / or the target video, and optimizing the coordinate information through the target area to obtain the actual object coordinates of the at least one object; mapping the actual object coordinates back to the system coordinates of the target image and / or the target video to obtain the actual positioning result of the at least one object in space.
[0006] Optionally, in an embodiment of the present application, the step of extracting feature description information of at least one object in a target image and / or a target video by using the multimodal large model, and generating initial structured data of the at least one object according to the feature description information includes: performing format conversion on an initial target image and / or an initial target video by using the multimodal large model to obtain the target image and / or the target video, and obtaining standardized format data containing semantic information and position data of the at least one object according to user requirements; based on the target image and / or the target video, recording change information of the at least one object to generate a time-series semantic description; based on the standardized format data, extracting the category, description, relative position information of the at least one object and its spatial association with surrounding items, so as to determine the initial structured data of the at least one object by combining user requirements and the time-series semantic description.
[0007] Optionally, in an embodiment of the present application, the step of generating structured data containing coordinate information and description information of the at least one object by combining the visual cues and the initial structured data includes: constructing a semantic map based on the initial structured data and the time-series semantic description; determining the structured data according to the semantic map and the visual cues.
[0008] Optionally, in an embodiment of the present application, the grid processing of the target image and / or the target video to add visual cues including position information to multiple grids in the target image and / or the target video includes: determining specification information of the grid processing based on user requirements, and performing grid processing on the target image and / or the target video according to the specification information; performing label processing on the multiple grids according to the mapping information between the multiple grids and the actual area to complete the addition of the visual cues.
[0009] Optionally, in an embodiment of the present application, the optimizing the coordinate information through the target area to obtain the actual object coordinates of the at least one object includes: generating a refined grid on the image or video frame corresponding to the target area; determining refined spatial coordinates of the at least one object in the refined grid to optimize the coordinate information and determine the actual object coordinates of the at least one object.
[0010] Optionally, in an embodiment of the present application, it further includes: determining whether the actual positioning result meets user requirements; if the actual positioning result meets user requirements, transmitting structured data corresponding to the actual positioning result to the target user by using the multimodal large model, otherwise correcting the actual positioning result according to the user requirements until it meets the user requirements.
[0011] An embodiment of the second aspect of the present application provides a device for enhancing the spatial perception ability of a multimodal large model, including: an extraction module, configured to use the multimodal large model to extract feature description information of at least one object in a target image and / or a target video, and generate initial structured data of the at least one object according to the feature description information; a processing module, configured to perform grid processing on the target image and / or the target video, and add visual cues including position information to multiple grids in the target image and / or the target video, so as to generate structured data including coordinate information and description information of the at least one object by combining the visual cues and the initial structured data; a positioning module, configured to, based on the structured data, locate a target area corresponding to the at least one object in the target image and / or the target video, and optimize the coordinate information through the target area to obtain the actual object coordinates of the at least one object; a mapping module, configured to map the actual object coordinates back to the system coordinates of the target image and / or the target video to obtain the actual positioning result of the at least one object in space.
[0012] Optionally, in an embodiment of the present application, the extraction module includes: a conversion unit configured to use the multimodal large model to perform format conversion on the initial target image and / or the initial target video to obtain the target image and / or the target video, and obtain standardized format data including semantic information and position data of the at least one object according to user requirements; a first generation unit configured to record change information of the at least one object based on the target image and / or the target video to generate a time-series semantic description; an extraction unit configured to extract the category, description, relative position information of the at least one object, and its spatial association with surrounding items based on the standardized format data, and determine initial structured data of the at least one object by combining user requirements and the time-series semantic description.
[0013] Optionally, in an embodiment of the present application, the processing module includes: a construction unit configured to construct a semantic map based on the initial structured data and the time-series semantic description; a determination unit configured to determine the structured data according to the semantic map and the visual cue.
[0014] Optionally, in an embodiment of the present application, the processing module includes: a first processing unit configured to determine specification information of the grid processing based on user requirements, and perform grid processing on the target image and / or the target video according to the specification information; a second processing unit configured to perform label processing on the plurality of grids according to mapping information between the plurality of grids and the actual area to complete the addition of the visual cue.
[0015] Optionally, in an embodiment of the present application, the positioning module includes: a second generation unit configured to generate a refined grid on an image or video frame corresponding to the target area; an optimization unit configured to determine refined spatial coordinates of the at least one object in the refined grid to optimize the coordinate information and determine actual object coordinates of the at least one object.
[0016] Optionally, in an embodiment of the present application, it further includes: a judgment module configured to judge whether the actual positioning result meets user requirements; a correction module configured to, when the actual positioning result meets user requirements, transmit structured data corresponding to the actual positioning result to the target user by using the multimodal large model, otherwise correct the actual positioning result according to user requirements until the user requirements are met.
[0017] An embodiment of the third aspect of the present application provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the program to implement the method for enhancing the spatial perception ability of the multimodal large model as described in the above embodiment.
[0018] In the fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the program is executed by a processor, the method for enhancing the spatial perception ability of the multi-modal large model as described above is implemented.
[0019] In the fifth aspect of the embodiments of the present application, a computer program product is provided, including a computer program, which is used to implement the method for enhancing the spatial perception ability of the multi-modal large model as described above when executed.
[0020] The embodiments of the present application can use a multi-modal large model to extract feature description information of objects in a target image and / or a target video and generate initial structured data; then perform grid processing on the target image and / or the target video and add visual cues containing position information to the grid to generate structured data containing the coordinate information and description information of the objects; then obtain the actual object coordinates of at least one object by performing refined grid processing on the target area corresponding to at least one object; finally, map the actual object coordinates back to the system coordinates of the target image and / or the target video to obtain the actual positioning result of at least one object in space. Thus, by combining multi-modal semantic extraction, explicit grid-based positioning, multi-frame memory mechanism, and semantic map construction, etc., the spatial understanding and precise positioning ability of the multi-modal large model in dynamic scenarios are improved, and target recognition and precise positioning from coarse-grained to fine-grained in complex scenarios are completed; moreover, the explicit grid-based positioning and semantic map construction in the present application can greatly reduce the complexity of the multi-modal large model in intelligent scene positioning tasks. At the same time, the present application supports diverse interaction and output forms, has extremely strong adaptability, effectively expands the application scenarios of the multi-modal large model, and can meet the requirements of various application scenarios. Thus, the problems in the related art are solved, where traditional vision methods and vision transformers limit the mapping ability and prediction ability of the precise coordinates of the multi-modal large model, resulting in deficiencies in the spatial positioning ability of the multi-modal large model, increasing the learning complexity of the multi-modal large model, making the multi-modal large model have poor adaptability to dynamic scenarios and limited application flexibility, and restricting its performance in actual precise space positioning tasks.
[0021] Additional aspects and advantages of the present application will be partly given in the following description, partly will become obvious from the following description, or will be understood through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The above and / or additional aspects and advantages of the present application will become obvious and easy to understand from the following description of the embodiments in conjunction with the drawings, where: Figure 1 is a flowchart of a method for enhancing the spatial perception ability of a multi-modal large model according to an embodiment of the present application; Figure 2Flow chart for dynamic update and global analysis of semantic map in an embodiment of the present application; Figure 3 Flow chart for method of enhancing spatial positioning ability of enhanced multi-modal large model in an embodiment of the present application; Figure 4 Block diagram of module composition for enhancing spatial perception ability of enhanced multi-modal large model in an embodiment of the present application; Figure 5 Flow chart for execution of object recognition and positioning based on image in an embodiment of the present application; Figure 6 Schematic diagram of execution process of first visual positioning function in an embodiment of the present application; Figure 7 Schematic diagram of execution process of second visual positioning function in an embodiment of the present application; Figure 8 Schematic diagram of result of example in the field of robot operation in an embodiment of the present application; Figure 9 Flow chart for visual question answering object positioning and operation based on user input in an embodiment of the present application; Figure 10 Schematic diagram of structure of device for enhancing spatial perception ability of enhanced multi-modal large model provided according to an embodiment of the present application; Figure 11 Schematic diagram of structure of electronic device provided according to an embodiment of the present application.
[0023] Reference numerals: 10 - Device for enhancing spatial perception ability of enhanced multi-modal large model: 100 - Extraction module, 200 - Processing module, 300 - Positioning module, 400 - Mapping module; 1101 - Memory, 1102 - Processor, 1103 - Communication interface. Detailed implementation manners
[0024] The embodiments of the present application are described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described by referring to the drawings are exemplary and are intended to explain the present application, and should not be construed as a limitation to the present application.
[0025] The method and device for enhancing the spatial perception ability of a multimodal large model according to the embodiments of the present application will be described below with reference to the accompanying drawings. In view of the problems in the related art mentioned in the above background art, that is, traditional vision methods and vision transformers limit the mapping ability and prediction ability of precise coordinates of multimodal large models, resulting in deficiencies in the spatial positioning ability of multimodal large models, increasing the learning complexity of multimodal large models, making multimodal large models have poor adaptability to dynamic scenes and limited application flexibility, and restricting their performance in actual precise spatial positioning tasks, the present application provides a method for enhancing the spatial perception ability of multimodal large models. In this method, feature description information of objects in a target image and / or a target video can be extracted by using a multimodal large model and initial structured data can be generated; then, the target image and / or the target video are subjected to grid processing and visual cues containing position information are added to the grids to generate structured data containing coordinate information and description information of the objects; then, the target regions corresponding to at least one object are subjected to refined grid processing to obtain the actual object coordinates of at least one object; finally, the actual object coordinates are mapped back to the system coordinates of the target image and / or the target video to obtain the actual positioning results of at least one object in space. Thus, by combining multimodal semantic extraction, explicit grid positioning, multi-frame memory mechanism, semantic map construction, etc., the spatial understanding and precise positioning ability of multimodal large models in dynamic scenes are improved, and target recognition and precise positioning from coarse-grained to fine-grained in complex scenes are completed; moreover, the explicit grid positioning and semantic map construction in the present application can greatly reduce the complexity of multimodal large models in intelligent scene positioning tasks. At the same time, the present application supports diverse interaction and output forms, has extremely strong adaptability, effectively expands the application scenarios of multimodal large models, and can meet the requirements of various application scenarios. Thus, the problems in the related art, that is, traditional vision methods and vision transformers limit the mapping ability and prediction ability of precise coordinates of multimodal large models, resulting in deficiencies in the spatial positioning ability of multimodal large models, increasing the learning complexity of multimodal large models, making multimodal large models have poor adaptability to dynamic scenes and limited application flexibility, and restricting their performance in actual precise spatial positioning tasks, are solved.
[0026] Specifically, Figure 1 is a flowchart of a method for enhancing the spatial perception ability of a multimodal large model provided by an embodiment of the present application.
[0027] As Figure 1 shown, the method for enhancing the spatial perception ability of the multimodal large model includes the following steps: In step S101, feature description information of at least one object in a target image and / or a target video is extracted by using a multimodal large model, and initial structured data of at least one object is generated according to the feature description information.
[0028] It can be understood that the target image and the target video can be understood here as the images and videos received and processed by the multimodal large model when performing precise spatial positioning tasks in practical applications. It can be, but is not limited to, the images and videos input by the user through the interactive terminal for multimodal interaction with the multimodal large model, such as natural language dialogue, voice input, image or video annotation, gesture commands, etc., and then processed by the multimodal large model itself.
[0029] In some embodiments, when enhancing the spatial perception ability of the multimodal large model, that is, enhancing the precise spatial positioning ability of the multimodal large model in practical applications, it can be, but is not limited to, first using the multimodal large model to extract the feature description information of at least one object in the target image or the target video or both the target image and the target video, so as to further accurately locate at least one object according to the feature description information.
[0030] It should be noted that at least one object can be understood here as the object that the multimodal large model needs to accurately locate. It is at least one, but can also be multiple. For example, the user inputs an image of a bookshelf and wants to find one book, two books, three books, etc. that they need from the image. Specifically, it can be determined by those skilled in the art or the user according to the actual application scenario and actual needs. The embodiments of the present application are only for illustrative purposes and are not specifically limited.
[0031] Moreover, the target image and / or the target video in the subsequent process can be understood as the target image or the target video or both the target image and the target video.
[0032] Furthermore, in order to reduce the processing complexity of the multimodal large model, the embodiments of the present application can also convert these feature description information into certain structured data for subsequent processing.
[0033] Optionally, in an embodiment of the present application, using the multimodal large model to extract the feature description information of at least one object in the target image and / or the target video, and generating the initial structured data of at least one object includes: using the multimodal large model to perform format conversion on the initial target image and / or the initial target video to obtain the target image and / or the target video, and obtaining the standardized format data including the semantic information and position data of at least one object according to the user's needs; based on the target image and / or the target video, recording the change information of at least one object to generate a time series semantic description; based on the standardized format data, extracting the category, description, relative position information of at least one object and its spatial association with surrounding items, and combining the user's needs and the time series semantic description to determine the initial structured data of at least one object.
[0034] It can be understood that the initial target image and / or the initial target video can be understood here as the unprocessed image or video or image and video input by the user.
[0035] During the actual execution process, when extracting the feature description information of at least one object, in order to enable the image or video or image and video input by the user to be processed quickly and efficiently in the multimodal large model, the present application can first perform format conversion on the initial target image and / or the initial target video to obtain the target image and / or the target video.
[0036] Specifically, the user can input the initial target image and / or the initial target video into the multimodal large model. After receiving it, the multimodal large model performs preprocessing on the target image and / or the target video, that is, uses image encoding technology to perform format conversion on the target image and / or the target video, including but not limited to steps such as image format conversion, noise removal, and contrast adjustment, to ensure that the image can adapt to subsequent processing, thereby providing a necessary basis for further analysis and positioning.
[0037] At the same time, considering that the scene requirements of users are different, the multimodal large model has a high processing ability for image information, but its text understanding ability may be insufficient. Therefore, the embodiment of the present application can also introduce a text processing module to extract the text in the image to determine the user's demand information, so as to obtain standardized format data including the semantic information and position data of at least one object according to the user's demand.
[0038] For example, the user inputs an image of a bookshelf and wants to find a book they need from the image. Except for the text on the book cover, there may already be text information such as "Please help me find the XX book" in the image. The introduced text processing module can extract the user's demand information. The standardized format data is: User demand: XXX; Target object name: XXX; Target object location: XXX, etc.; This standardized format data can provide a reference for the output format of structured data in the subsequent process.
[0039] Thus, the embodiment of the present application can realize the joint processing of the user's demand and image data using a standardized process, so as to generate standardized format data that meets the standardized requirements and includes the semantic information and position data of at least one object as the input for the subsequent steps.
[0040] Next, embodiments of the present application can record the item change information of at least one object in a dynamic scene based on the target image and / or target video, in combination with a multi-frame memory mechanism, and generate a time-series semantic description to adapt to the requirements of the dynamic environment. That is, embodiments of the present application can record the change information of at least one object in the scene by combining multi-frame image data, thereby determining the position of the object, and effectively reducing the positioning error. For example, if object A is on the desk in 99 out of 100 frames of image data and on the chair in 1 frame of image data, it can be determined that the position of object A has changed and is located at different positions at different times.
[0041] Finally, embodiments of the present application can, based on the format of the standardized format data, use a multimodal large model to perform semantic analysis on the image by combining visual and language modalities. Through the combined analysis of the visual features and language modality of the object, the model can identify and extract the feature description information of at least one object in the image, including but not limited to the category, description, spatial association between the object and surrounding items, and the relative position relationship between the two. Furthermore, by combining user requirements and the time-series semantic description, initial structured data containing the semantic information and position data of at least one object can be generated.
[0042] Among them, semantic information refers to the category, description, etc. in the feature description information; position data refers to the position of the item preliminarily calibrated based on the explicit spatial relative position reference, according to the spatial association between the object and surrounding items and the relative position relationship between the two, such as object A is on top of the desk, etc. And, the initial structured data can be simply understood here as structured item information description data, which includes but is not limited to the semantic information and position data of at least one object, etc.
[0043] Embodiments of the present application can convert the original image information into a standardized format and generate initial structured data containing the semantic information and position data of at least one object, providing an important basis for subsequent positioning and recognition tasks.
[0044] Step S102, perform grid processing on the target image and / or target video, and add visual cues containing position information to multiple grids of the target image and / or target video, so as to generate structured data containing the coordinate information and description information of at least one object by combining the visual cues and the initial structured data.
[0045] In some embodiments, the present application can perform grid processing on the target image and / or target video, so as to add visual cues containing position information to multiple grids of the target image and / or target video, and then use a multimodal large model to generate structured data containing object coordinate information and description information.
[0046] Among them, adding visual cues containing location information to multiple grids of the target image and / or target video can be understood as adding visual elements or markers that can represent the position coordinates of the grid in the image to the grids that have been divided in the target image and / or target video.
[0047] Then, by combining this visual cue with the initial structured data that only contained the semantic information and preliminary position data of at least one object before, structured data containing the coordinate information and description information of at least one object can be generated. Among them, the coordinate information can be understood here as the spatial coordinates of at least one object in the target image and / or target video obtained according to the visual cues of the grid after adding visual cues containing location information to multiple grids of the target image and / or target video, and the description information includes all the information in the initial structured data.
[0048] And the structured data containing the coordinate information and description information of at least one object can be understood here as structured data that adds more accurate coordinate information of at least one object in the grids of the target image and / or target video on the basis of the initial structured data containing semantic information and position data. This data combines the attributes and spatial relationships of the object, which helps to improve the recognition accuracy of the object.
[0049] Next, the processes of grid processing the target image and / or target video in the embodiments of the present application to add visual cues containing location information to multiple grids of the target image and / or target video, and combining the visual cues and the initial structured data to generate structured data containing the coordinate information and description information of at least one object will be further described.
[0050] Optionally, in an embodiment of the present application, grid processing the target image and / or target video to add visual cues containing location information to multiple grids of the target image and / or target video includes: determining the specification information of the grid processing based on user requirements, and performing grid processing on the target image and / or target video according to the specification information; performing label processing on the multiple grids according to the mapping information between the multiple grids and the actual area to complete the addition of the visual cues.
[0051] In some embodiments, when grid processing the target image and / or target video, the target image and / or target video is divided into multiple grids according to a fixed size. Considering that users have different requirements for accuracy and response speed, the present application can determine the specification information of the grid processing according to user requirements, and then dynamically adjust the grid according to the specification information.
[0052] Among them, the specification information of grid processing includes but is not limited to grid size, numbering rules or positioning accuracy, so as to ensure that the grid processing in the embodiment of the present application can adapt to complex scenarios and different task requirements.
[0053] Furthermore, the embodiment of the present application can label multiple grids according to the mapping information between multiple grids and actual areas, that is, unique numbers or coordinates are marked for each grid. For example, a certain image is divided into 100*100 grids, and the target area corresponding to the position of at least one object in the actual area is the grid where the 66th column from left to right and the 66th row from bottom to top intersect in the grid, and the label information can be 66′66, etc. In this way, visual cues containing location information are added to multiple grids of the target image and / or target video.
[0054] It should be noted that professionals in this technical field can determine the type of visual prompts according to actual application scenarios and actual needs, including but not limited to bounding boxes, labels, pixel-level prompts, and soft prompts. These visual prompts can provide additional information to guide the model to interpret and process visual data, thereby enhancing the model's visual perception ability.
[0055] The embodiment of the present application can generate visual cues containing location information in the grid to ensure that each grid is associated with the location information of the area where it is located, and at the same time make it easier for the multimodal large model to understand these coded information, thereby improving the multimodal large model's ability to predict the precise coordinates of objects, thereby improving the performance of the multimodal large model in complex understanding and reasoning tasks. It can also provide an identification basis for precise positioning, reduce the learning complexity of the multimodal large model, and facilitate object positioning and subsequent precise labeling.
[0056] Optionally, in one embodiment of the present application, structured data containing coordinate information and description information of at least one object is generated by combining visual cues and initial structured data, including: constructing a semantic map based on the initial structured data and time series semantic description; and determining structured data based on the semantic map and visual cues.
[0057] In other embodiments, after the target image and / or target video is gridded and visual cues containing location information are added to the grid, the present application can input the gridded image into a multimodal large model, and the multimodal large model can construct a semantic map based on the initial structured data and the time series semantic description, and then determine the structured data of at least one object based on the semantic map and the visual cues. Thus, it is possible to generate structured data containing the coordinate information and description information of at least one object by combining the visual cues and the initial structured data.
[0058] For example, this application can combine multi-frame image data to record the dynamic changes of an object over time, including continuous changes in position, state, and description. Then, based on the structured data and time series information, a semantic map is constructed to achieve the global semantic data integration of multi-frame images.
[0059] Among them, the semantic map here refers to the dynamic structured data constructed by combining multi-frame image data, the spatial position information of the object, and the time series semantic description, which is a form of expression of the structured data in the embodiments of this application. It can record and reflect in real time the position changes, state changes of the object in space, and the spatial relationship with surrounding objects. For example, the semantic map records the preliminary positioning results of the object in the form of a JSON file, including the item name, horizontal and vertical coordinate information, timestamp, and the spatial relationship with surrounding objects.
[0060] This semantic map not only contains the semantic information and position information of the object, but also can store the spatial association between the object and other items, support the maintenance of continuous time series data, and adapt to complex and changing environments. Through this semantic map, the embodiments of this application can maintain the association between historical information and current frame data, ensure the global understanding ability of the multi-modal large model in dynamic scenarios, and ensure the precise positioning ability of the multi-modal large model in dynamic scenarios.
[0061] Furthermore, inputting the semantic map and the current frame image data into the multi-modal large model together is more conducive to the joint analysis of the multi-modal large model, so as to generate more accurate item position coordinate information and semantic data.
[0062] It should be noted that the semantic map in the embodiments of this application is dynamically changing, rather than fixed. For example, Figure 2 is the flowchart of the dynamic update and global analysis of the semantic map for an embodiment of this application. As Figure 2 shown, after receiving the target video, the embodiments of this application will process the multi-frame images in the video, use the multi-frame memory mechanism to integrate the information of the front and back frames, and generate a time series semantic map. Or, after receiving multiple target images with a time series, the embodiments of this application can also use the multi-frame memory mechanism to integrate the information of the front and back frames to generate a time series semantic map.
[0063] After generating the time series semantic map, the embodiments of this application will judge whether the scene has changed as the number of multi-frame images increases. If the scene has not changed, the semantic map information will be kept unchanged, and the semantic map and the current frame image will be input into the multi-modal large model for multi-modal analysis to generate the structured data of at least one object. This data includes but is not limited to the description, coordinates, and spatial relationship information of the object for subsequent processing.
[0064] If the scene changes, the items that appear in the current frame and their dynamic positions are added according to the scene change information, or the positions or states of the objects in the historical frames are modified, or the items that no longer exist in the semantic map are deleted, so as to ensure that the semantic map can reflect the changes of objects in the dynamic environment in real time.
[0065] It should be noted that when generating the initial structured data, in order to initially determine the positions of at least one object, the multimodal large model in the embodiments of the present application can also construct a semantic map by combining multi-frame image data. However, the semantic map constructed by the initial structured data is for initially determining the positions of the objects, and the semantic map constructed by the structured data is for further determining the more accurate positions of the objects. The modification time of the semantic map is not specifically limited in the embodiments of the present application and can be achieved during the generation of both the initial structured data and the structured data. It can be set or adjusted by those skilled in the art according to the actual situation and actual needs.
[0066] Step S103: Based on the structured data, locate the target regions corresponding to at least one object in the target image and / or target video, and optimize the coordinate information through the target regions to obtain the actual object coordinates of at least one object.
[0067] As a possible implementation, after performing grid processing on the target image and / or target video and adding visual cues containing position information to the grids to obtain structured data including the coordinate information and description information of at least one object, in order to further improve the accuracy of the spatial coordinates of at least one object, the embodiments of the present application can also re-locate the target regions corresponding to at least one object in the target image and / or target video based on the structured data, and then further refine the spatial coordinates of at least one object within the target regions, thereby optimizing the previous spatial coordinates to obtain the actual object coordinates of at least one object.
[0068] Herein, the target region refers to the region corresponding to the grid position where at least one object is located after the grid division of the target image and / or target video. For example, when a certain image is divided into 100*100 grids, the target region corresponding to at least one object is the grid at the intersection of the 66th column from left to right and the 66th row from bottom to top, with the label information of 66′66.
[0069] It should be noted that before further processing the target area, the embodiments of the present application will first determine whether the target area meets the positioning requirements, that is, determine whether there is at least one object to be located in the target area, or whether the target area meets the positioning requirements of at least one object, such as maintaining the integrity of the object within the target area. When the target area meets the positioning requirements, the system will continue with subsequent processing. If the target area does not meet the recognition requirements, the embodiments of the present application can adjust the target area based on the preliminary detection results and re - conduct area analysis to ensure the accuracy and efficiency of the final recognition.
[0070] Optionally, in an embodiment of the present application, the actual object coordinates of at least one object are obtained by optimizing the coordinate information of the target area, including: generating a refined grid on the image or video frame corresponding to the target area; determining the refined spatial coordinates of at least one object in the refined grid to optimize the coordinate information and determine the actual object coordinates of at least one object.
[0071] In some embodiments, when the present application optimizes the coordinate information of at least one object through the target area to obtain the actual object coordinates of at least one object, it can, but is not limited to, accurately locate the target area corresponding to at least one object in the target image and / or target video according to the name and coordinates of the object in the structured data, then perform cropping, magnification, and refined grid processing on the target area to generate a refined grid, and then determine the refined spatial coordinates of at least one object in the refined grid to optimize the previous coordinate information, that is, take this refined spatial coordinate as the latest spatial coordinate, thereby determining the actual object coordinates of at least one object.
[0072] For example, the present application can crop the target area in the target image and / or target video according to the previous spatial coordinate information, that is, take the area corresponding to the grid where at least one object is located as an image whole, crop out this image whole from the target image and / or target video and generate a refined grid on this image whole.
[0073] Then, the embodiments of the present application can call a multimodal large - model to combine visual and language modalities to further extract the precise position, boundary information of the target item, and the spatial association with surrounding items; combine semantic maps and time - series information to analyze the dynamic changes of the items, optimize the positioning accuracy of the target area, and determine the refined spatial coordinates of at least one object.
[0074] Finally, the embodiments of the present application can generate precise structured data including object category, description, location, and dynamic relationship based on this refined spatial coordinate, realizing the high - precision positioning requirements in complex scenarios, especially dynamic scenarios.
[0075] It should be noted that the process of locating the target area corresponding to at least one object in the embodiments of the present application and then further refining the spatial coordinates of at least one object within the target area to optimize the previous spatial coordinates is not a one-time process, but can be iteratively optimized according to the actual needs and actual situations of those skilled in the art and users until the accuracy requirements of the spatial positioning task or user requirements are met.
[0076] Step S104: Map the actual object coordinates back to the system coordinates of the target image and / or target video to obtain the actual positioning result of at least one object in space.
[0077] During actual execution, after obtaining the actual object coordinates of at least one object, the present application can map the actual object coordinates of the local image back to the system coordinates of the original target image and / or original target video to obtain the actual positioning result of at least one object in the space of the original target image and / or target video.
[0078] During this process, for the convenience of the output of the multimodal large model and user understanding, the embodiments of the present application can output the actual positioning result to the user in the form of structured data. Among them, the actual positioning result includes, but is not limited to, the description, coordinates, and spatial relationship information of the object.
[0079] Specifically, the embodiments of the present application can map the precise position of the cropped area back to the global coordinate system, generate semantic data with dynamically changing labels in combination with the time series; and output the semantic information of the item through the natural language generation module to clearly describe the category, position, and status of the target object; finally, generate operation instructions or non-text tokens (such as action tokens) corresponding to the user's needs and output them to the user.
[0080] Among them, action tokens are a specific type of non-text tokens used in the field of robotics to represent the specific actions or tasks that a robot should perform. By combining semantic map data, it can output the result in the form of natural language description, operation instructions, or precise coordinates, thereby improving the positioning accuracy and output flexibility of objects in complex scenarios.
[0081] Optionally, in an embodiment of the present application, it further includes: determining whether the actual positioning result meets the user's needs; if the actual positioning result meets the user's needs, use the multimodal large model to transmit the structured data corresponding to the actual positioning result to the target user, otherwise correct the actual positioning result according to the user's needs until it meets the user's needs.
[0082] In some embodiments, after generating the actual positioning result, the present application will perform self-evaluation and detection on the generated result. For example, it will perform verification in combination with confidence thresholds, historical frame records, etc. If a large deviation from the expectation is found, secondary positioning or multi-frame correction will be triggered.
[0083] Furthermore, the embodiments of the present application will also determine whether the output actual positioning result meets the user's requirements. If it meets the requirements, the structured data corresponding to the actual positioning result will be transmitted to the target user using the multi-modal large model, and the processing will end; if it does not meet the requirements, adjustments will be made according to the analysis results to ensure that the final output meets the user's requirements.
[0084] For example, the user inputs a 20s target video and requests to find the video time period when the book appears in the video. After receiving the positioning task and the target video, the multi-modal large model finds that a corner of the book appears in the upper right corner of the video in the first 0.5s, the book does not appear between 0.5s and 3s, and the complete book appears in the video from 3s to 20s. So the output positioning result is from 3s to 20s of the video, and the book is located at the intersection of the 66th column grid from left to right and the 66th row grid from bottom to top after the video is divided into 100*100 grids. However, what the user needs is all the positioning results when the book appears, rather than the positioning results when the complete book appears. At this time, adjustments need to be made according to the analysis results to ensure that the final output meets the user's requirements.
[0085] Among them, the structured data output by the embodiments of the present application can be presented in various forms, such as natural language descriptions, precise coordinates, operation instructions, or behavior instructions based on special marks, etc., providing rich interaction methods for users and adapting to the diverse application requirements in complex scenarios at the same time.
[0086] The following uses a specific embodiment to explain the present application in detail.
[0087] Figure 3 It is a flowchart of a method for enhancing the spatial positioning ability of the multi-modal large model according to an embodiment of the present application; Figure 4 It is a schematic diagram of the method principle for enhancing the spatial perception ability of the multi-modal large model according to an embodiment of the present application. Among them, each module in the system can also be independently embedded in the multi-modal large model. Here, only an exemplary illustration is made and no specific limitation is imposed. As Figure 3 and Figure 4 shown: Step S301, using the image information extraction function, perform preliminary processing on the input image or video, extract the feature description information of the object, and generate initial structured data, including the category, description of the object, and the relative position relationship of the object, etc.; Step S302: Using the first visual positioning function, perform grid processing on the input image or video, generate a preliminary recognition result containing the object position and description information, and convert it into structured data containing the object's spatial coordinates (coordinate information) and description information. Step S303: Using the second visual positioning function, combine the object name and coordinates (the position of the item recognized by the first visual positioning function), accurately locate the target area in the image, perform cropping, magnification, and refined grid processing on the target area, and use the system to further optimize the object's precise spatial coordinates. Step S304: Map the fine coordinates of the local image back to the system coordinates of the original image, generate structured data containing the accurate object positioning result, and according to the user's needs, output it in various forms (such as natural language description, accurate coordinates, operation instructions, or behavior instructions based on special marks) through the result output function, which is convenient for subsequent analysis and application, and at the same time adapts to the diverse application requirements in complex scenarios.
[0088] Next, the embodiments of the present application combine Figure 5 、 Figure 6 and Figure 7 to further illustrate the specific execution process of steps S201 - S203. Among them, Figure 5 is the execution flowchart of image - based target recognition and positioning according to an embodiment of the present application, Figure 6 is the schematic execution flow diagram of the first visual positioning function according to an embodiment of the present application, Figure 7 is the schematic execution flow diagram of the second visual positioning function according to an embodiment of the present application, as shown in Figure 5 、 Figure 6 and Figure 7 shown: First, pre - process the picture to generate standardized and formatted data containing semantic information and position data; then use a multi - modal large - model for semantic analysis to extract the feature description information of the object to generate initial structured data, that is, structured object information description data. Then, as shown in Figure 6 , perform grid processing on the image through the first visual positioning function, add a unique number or coordinate information to each grid to help locate the relative position of the object; on this basis, according to the semantic information of the object, call the multi - modal large - model to identify the object, perform a coarse - grained spatial positioning of the object's position, generate a coarse - grained recognition result of structured output data containing the object position and description information, and further support fine - grained image positioning.
[0089] Next, as shown in Figure 7As shown, through the second visual positioning function, the rough-grained recognition results of the image are further processed to achieve precise object positioning. First, using the rough-grained image recognition results, the description and position information of a single object are extracted. Based on this information, the operation of segmenting the sub-graph of the target object is performed in the original image for further processing. On this basis, through grid processing, the image is divided into multiple grids, and coding or coordinate information is added to each grid to assist subsequent precise positioning. Next, a multi-modal large model is called to recognize the gridified image and generate fine-grained object positioning results. This process ensures that the precise position of the target object can be accurately recognized and provides accurate data support for subsequent tasks.
[0090] And, Figure 8 is a schematic diagram of the results of an example in the field of robot operation for an embodiment of the present application. As Figure 8 shown, the embodiment of the present application can achieve dynamic scene target recognition and precise grasping through a multi-modal large model. Among them, the multi-modal large model can use, for example, GPT-4o or LLaMA 3.2 Vision, etc., which can be specifically selected or adjusted by professionals in the field according to the actual situation. Here, only an exemplary illustration is given and no specific limitation is made.
[0091] Specifically, the user can send a demand to the system through a natural language instruction, such as "I'm a bit thirsty". After receiving the instruction, the multi-modal large model performs semantic parsing on the instruction, extracts the clear operation demand from it, such as "grasp the water cup". On the basis of semantic parsing, the system calls the visual sensor to obtain the image of the current scene, and performs preliminary processing on the image, extracts the basic feature description information of the object, and generates structured data, including the category, description of the object, and the relative position relationship of the object, etc.
[0092] Then, the multi-modal large model performs rough-grained recognition on the target object in the scene. By performing grid processing on the image, unique coordinate information is marked for each grid, and based on the visual features and semantic information of the object, the category, description of the target object, and its rough position in the scene are recognized. A semantic map is generated based on the recognition results.
[0093] Based on coarse-grained recognition, the system further crops the region where the target object is located and performs a refined grid process on the cropped region to improve the positioning accuracy of the target object. At this time, the semantic map provides context information and global spatial association information to assist the system in more efficiently and accurately analyzing the cropped region. The multi-modal large model is used to further process the cropped image to identify the boundary features, shape of the target object, and its spatial relationship with the surrounding environment. Combining the semantic map and the detailed analysis results of the cropped region, the system can generate the precise position coordinates of the target object and map these coordinates to the global coordinate system of the original image. The generated output data includes the precise coordinates of the target object, dynamic semantic descriptions, and spatial relationships with other objects, such as "the water cup is located at the center of the table, close to the left side of the book".
[0094] The semantic map not only plays a role in the initial recognition stage but also can dynamically update the time series information of objects, such as recording the movement trajectory or position change of the object. This can ensure that even when the scene changes, the system can accurately capture the state of the target object and maintain real-time global semantic understanding.
[0095] According to the precise position of the target object and the user's needs, the system generates corresponding operation instructions and sends the instructions to the robot execution module. For example, the instructions include the coordinates of the target object and specific action descriptions, such as "grab the water cup at position (x, y, z)". After receiving the instructions, the robot uses the robotic arm to perform positioning and grasping operations and delivers the water cup to the user's hand or a specified position. After the operation is completed, the system will update the semantic map and generate a task execution report, including task execution status, time used, and final position confirmation information, for further analysis by the user or the system.
[0096] Thus, by combining the advantages of the multi-modal large model in visual semantic understanding and spatial positioning, through the integration of natural language interaction and robot dynamic operation, precise object recognition and efficient grasping operations in complex scenarios are achieved, significantly enhancing the scene adaptation ability of the robot in a dynamic environment. It should be noted that the method for enhancing the spatial perception ability of the multi-modal large model in this application is not limited to the field of robot operation and can also be widely applied to fields such as automatic control of smart homes, environment perception and path planning in autonomous driving, target item management and grasping in intelligent warehousing systems, and precise detection of lesion areas in medical images. And the specific implementation methods in different fields can be adjusted by professional technical personnel in this field according to the actual application scenarios and actual needs to achieve a comprehensive improvement in the spatial positioning ability of this application and meet diverse application requirements at the same time.
[0097] Furthermore, the embodiments of this application will make certain corrections to the output results. Figure 9The flowchart of object localization and operation for visual question answering (VQA) based on user input according to an embodiment of the present application. As Figure 9 shown, after obtaining the object description information from the text description provided by the user in the embodiment of the present application, it will be determined whether the object information exists in the semantic information. If the object information exists in the semantic information, the object localization will continue based on this information; if the object information does not exist, the first visual localization function will be called for coarse-grained object recognition and the approximate position of the object will be extracted. After the coarse-grained recognition, the second visual localization function will be called for fine-grained object recognition to further accurately locate the position of the target object. The fine-grained recognition module extracts the accurate position of the object through high-resolution image processing and precise grid division technology. Finally, according to the recognition result, the position or operation instruction of the object is output. The output result can be used to control the device to perform precise operations on the target object, such as localization, grasping, etc.
[0098] According to the method for enhancing the spatial perception ability of the multi-modal large model proposed in the embodiment of the present application, the feature description information of the object in the target image and / or target video can be extracted by the multi-modal large model and the initial structured data can be generated; then the target image and / or target video are meshed and visual cues containing position information are added to the grid, and the structured data containing the coordinate information and description information of the object are generated in combination with the visual cues; then the actual object coordinates of at least one object are obtained by performing refined grid processing on the target area corresponding to at least one object; finally, the actual object coordinates are mapped back to the system coordinates of the target image and / or target video to obtain the actual localization result of at least one object in space. Thus, by combining multi-modal semantic extraction, explicit grid localization, multi-frame memory mechanism and semantic map construction, etc., the spatial understanding and precise localization ability of the multi-modal large model in dynamic scenes are improved, and the target recognition and precise localization from coarse-grained to fine-grained in complex scenes are completed; and the explicit grid localization and semantic map construction in the present application can greatly reduce the complexity of the multi-modal large model in the intelligent scene localization task. At the same time, the present application supports diverse interaction and output forms, has extremely strong adaptability, effectively expands the application scenarios of the multi-modal large model, and can meet the requirements of various application scenarios. Thus, it solves the problems in the related art that traditional vision methods and vision transformers limit the mapping ability and prediction ability of the precise coordinates of the multi-modal large model, resulting in deficiencies in the spatial localization ability of the multi-modal large model, increasing the learning complexity of the multi-modal large model, making the multi-modal large model have poor adaptability to dynamic scenes and limited application flexibility, and restricting its performance in actual precise spatial localization tasks.
[0099] Next, an apparatus for enhancing the spatial perception ability of the multi-modal large model according to an embodiment of the present application will be described with reference to the accompanying drawings.
[0100] Figure 10 It is a schematic structural diagram of a device for enhancing the spatial perception ability of a multi-modal large model according to an embodiment of the present application.
[0101] As Figure 10 shown, the device 10 for enhancing the spatial perception ability of the multi-modal large model includes: an extraction module 100, a processing module 200, a positioning module 300, and a mapping module 400.
[0102] Among them, the extraction module 100 is used to extract the feature description information of at least one object in the target image and / or target video by using the multi-modal large model, and generate the initial structured data of at least one object according to the feature description information.
[0103] The processing module 200 is used to perform grid processing on the target image and / or target video, and add visual cues containing position information to multiple grids of the target image and / or target video, so as to generate structured data containing the coordinate information and description information of at least one object by combining the visual cues and the initial structured data.
[0104] The positioning module 300 is used to locate the target area corresponding to at least one object in the target image and / or target video based on the structured data, and optimize the coordinate information through the target area to obtain the actual object coordinates of at least one object.
[0105] The mapping module 400 is used to map the actual object coordinates back to the system coordinates of the target image and / or target video to obtain the actual positioning result of at least one object in space.
[0106] Optionally, in an embodiment of the present application, the extraction module 100 includes: a conversion unit, a first generation unit, and an extraction unit.
[0107] Among them, the conversion unit is used to perform format conversion on the target image and / or target video to obtain standardized format data containing the semantic information and position data of at least one object in combination with user requirements.
[0108] The first generation unit is used to record the change information of at least one object based on the target image and / or target video to generate a time series semantic description.
[0109] The extraction unit is used to extract the category, description, relative position information of at least one object and its spatial association with surrounding items based on the standardized format data, and determine the initial structured data of at least one object in combination with user requirements and the time series semantic description.
[0110] Optionally, in an embodiment of the present application, the processing module 200 includes: a construction unit and a determination unit.
[0111] Among them, the construction unit is used to construct a semantic map based on the initial structured data and the time-series semantic description.
[0112] The determination unit is used to determine the structured data according to the semantic map and the visual cue.
[0113] Optionally, in an embodiment of the present application, the processing module 200 includes: a first processing unit and a second processing unit.
[0114] Among them, the first processing unit is used to determine the specification information of grid processing based on the user's needs, so as to perform grid processing on the target image and / or target video according to the specification information.
[0115] The second processing unit is used to perform label processing on multiple grids according to the mapping information between the multiple grids and the actual area, so as to complete the addition of visual cues.
[0116] Optionally, in an embodiment of the present application, the positioning module 300 includes: a second generation unit and an optimization unit.
[0117] Among them, the second generation unit is used to generate refined grids on the image or video frame corresponding to the target area.
[0118] The optimization unit is used to determine the refined spatial coordinates of at least one object in the refined grid, so as to optimize the coordinate information and determine the actual object coordinates of at least one object.
[0119] Optionally, in an embodiment of the present application, it further includes: a judgment module and a correction module.
[0120] Among them, the judgment module is used to judge whether the actual positioning result meets the user's needs.
[0121] The correction module is used to transmit the structured data corresponding to the actual positioning result to the target user by using the multimodal large model when the actual positioning result meets the user's needs, otherwise correct the actual positioning result according to the user's needs until it meets the user's needs.
[0122] It should be noted that the foregoing explanation of the method embodiment for enhancing the spatial perception ability of the multimodal large model also applies to the device for enhancing the spatial perception ability of the multimodal large model in this embodiment, and will not be elaborated here.
[0123] The device for enhancing the spatial perception ability of a multimodal large model according to an embodiment of the present application can use the multimodal large model to extract feature description information of objects in a target image and / or target video and generate initial structured data; then perform grid processing on the target image and / or target video and add visual cues containing position information to the grid to generate structured data containing the coordinate information and description information of the objects; further perform refined grid processing on the target area corresponding to at least one object to obtain the actual object coordinates of at least one object; and finally map the actual object coordinates back to the system coordinates of the target image and / or target video to obtain the actual positioning results of at least one object in space. Thus, by combining multimodal semantic extraction, explicit grid-based positioning, multi-frame memory mechanism, and semantic map construction, etc., the spatial understanding and precise positioning ability of the multimodal large model in dynamic scenarios are improved, and target recognition and precise positioning from coarse-grained to fine-grained in complex scenarios are completed; moreover, the explicit grid-based positioning and semantic map construction in the present application can greatly reduce the complexity of the multimodal large model in intelligent scene positioning tasks. At the same time, the present application supports diverse interaction and output forms, has extremely strong adaptability, effectively expands the application scenarios of the multimodal large model, and can meet the requirements of various application scenarios. Thus, it solves the problems in the related technologies that traditional vision methods and vision transformers limit the mapping ability and prediction ability of precise coordinates of the multimodal large model, resulting in deficiencies in the spatial positioning ability of the multimodal large model, increasing the learning complexity of the multimodal large model, making the multimodal large model have poor adaptability to dynamic scenarios and limited application flexibility, and restricting its performance in actual precise space positioning tasks.
[0124] Figure 11 The following is a schematic structural diagram of an electronic device provided by an embodiment of the present application. The electronic device may include: A memory 1101, a processor 1102, and a computer program stored on the memory 1101 and executable on the processor 1102.
[0125] When the processor 1102 executes the program, it implements the method for enhancing the spatial perception ability of the multimodal large model provided in the above embodiment.
[0126] Furthermore, the electronic device further includes: A communication interface 1103 for communication between the memory 1101 and the processor 1102.
[0127] The memory 1101 is used to store a computer program executable on the processor 1102.
[0128] The memory 1101 may include a high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory.
[0129] If the memory 1101, the processor 1102, and the communication interface 1103 are implemented independently, the communication interface 1103, the memory 1101, and the processor 1102 can be interconnected through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity in representation, Figure 11 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0130] Optionally, in a specific implementation, if the memory 1101, the processor 1102, and the communication interface 1103 are integrated on a single chip, the memory 1101, the processor 1102, and the communication interface 1103 can communicate with each other through an internal interface.
[0131] The processor 1102 may be a Central Processing Unit (CPU), or an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0132] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the method for enhancing the spatial perception ability of the enhanced multimodal large model as described above is implemented.
[0133] The embodiments of the present application also provide a computer program product, including a computer program, and the computer program can run computer instructions, and when the computer instructions are executed by a processor, the method for enhancing the spatial perception ability of the enhanced multimodal large model provided by the embodiments of the present application is implemented.
[0134] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0135] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of this application, the meaning of "N" is at least two, such as two, three, etc., unless otherwise specifically defined.
[0136] Any process or method description depicted in a flowchart or described in other ways herein can be understood to represent a module, segment, or portion of code including one or N executable instructions for implementing a customized logical function or process, and the scope of the preferred embodiments of this application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of this application pertain.
[0137] The logic and / or steps represented in the flowchart or otherwise described herein can, for example, be considered a definitional sequence list of executable instructions for implementing logical functions, which can be embodied specifically in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in conjunction with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion (electronic device) having one or N wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, as the program can be obtained electronically by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.
[0138] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having suitable combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), and the like.
[0139] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of the above-described embodiments can be completed by a program instructing relevant hardware, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0140] In addition, each functional unit in various embodiments of the present application may be integrated into a processing module, may exist separately as individual physical units, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0141] The above-mentioned storage medium may be a read-only memory, a magnetic disk or an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A method for enhancing the spatial perception ability of a multimodal large model, characterized in that, Including the following steps: Using a multimodal large model to extract the feature description information of at least one object in the target image and / or target video, and generating the initial structured data of the at least one object according to the feature description information; Performing grid processing on the target image and / or the target video, and adding visual cues containing position information to multiple grids of the target image and / or the target video, so as to generate structured data containing the coordinate information and description information of the at least one object by combining the visual cues and the initial structured data; Based on the structured data, locating the target area corresponding to the at least one object in the target image and / or the target video, and optimizing the coordinate information through the target area to obtain the actual object coordinates of the at least one object; Mapping the actual object coordinates back to the system coordinates of the target image and / or the target video to obtain the actual positioning result of the at least one object in space.
2. The method for enhancing the spatial perception ability of the enhanced multi-modal large model according to claim 1, wherein, The step of using a multimodal large model to extract the feature description information of at least one object in the target image and / or target video, and generating the initial structured data of the at least one object according to the feature description information includes: Using the multimodal large model to perform format conversion on the initial target image and / or initial target video to obtain the target image and / or target video, and obtaining standardized format data containing the semantic information and position data of the at least one object according to user requirements; Based on the target image and / or target video, recording the change information of the at least one object to generate a time series semantic description; Based on the standardized format data, extracting the category, description, relative position information of the at least one object and its spatial association with surrounding items, and determining the initial structured data of the at least one object by combining user requirements and the time series semantic description.
3. The method for enhancing the spatial perception ability of the multi-modal large model according to claim 2, wherein, The step of generating structured data containing the coordinate information and description information of the at least one object by combining the visual cues and the initial structured data includes: Constructing a semantic map based on the initial structured data and the time series semantic description; Determining the structured data according to the semantic map and the visual cues.
4. The method for enhancing the spatial perception ability of the multi-modal large model according to claim 1, characterized in that, The step of performing grid processing on the target image and / or the target video to add visual cues containing position information to multiple grids of the target image and / or the target video includes: Determining the specification information of the grid processing according to user requirements, and performing grid processing on the target image and / or the target video according to the specification information; Performing label processing on the multiple grids according to the mapping information between the multiple grids and the actual area to complete the addition of the visual cues.
5. The method for enhancing the spatial perception ability of the enhanced multi-modal large model according to claim 1, characterized in that The step of optimizing the coordinate information through the target area to obtain the actual object coordinates of the at least one object includes: Generating a refined grid on the image or video frame corresponding to the target area; Determining the refined spatial coordinates of the at least one object in the refined grid to optimize the coordinate information and determine the actual object coordinates of the at least one object.
6. The method for enhancing the spatial perception ability of the enhanced multi-modal large model according to claim 1, wherein, It further includes: judging whether the actual positioning result meets the user's requirements; if the actual positioning result meets the user's requirements, transmitting the structured data corresponding to the actual positioning result to the target user by using the multi-modal large model; otherwise, correcting the actual positioning result according to the user's requirements until it meets the user's requirements.
7. An apparatus for enhancing the spatial perception ability of a multimodal large model, characterized in that, It includes: an extraction module, configured to extract the feature description information of at least one object in the target image and / or target video by using a multi-modal large model, and generate the initial structured data of the at least one object according to the feature description information; a processing module, configured to perform grid processing on the target image and / or the target video, and add visual cues containing position information to multiple grids of the target image and / or the target video, so as to generate structured data containing the coordinate information and description information of the at least one object by combining the visual cues and the initial structured data; a positioning module, configured to locate the target area corresponding to the at least one object in the target image and / or the target video based on the structured data, and optimize the coordinate information through the target area to obtain the actual object coordinates of the at least one object; a mapping module, configured to map the actual object coordinates back to the system coordinates of the target image and / or the target video to obtain the actual positioning result of the at least one object in space.
8. An electronic device, characterized in that, It includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the program to implement the method for enhancing the spatial perception ability of the multi-modal large model according to any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to be used for implementing the method for enhancing the spatial perception ability of the multi-modal large model according to any one of claims 1-6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed, it is used for implementing the method for enhancing the spatial perception ability of the multi-modal large model according to any one of claims 1-6.
Citation Information
Patent Citations
Semantic vision positioning method and device based on multi-modal graph convolutional network
CN111783457A
Defect labeling method based on computer vision large model
CN116912830A
Method, model and device for enhancing visual perception capability of multi-modal large language model
CN118585954A
Using grounded rationales to improve visual reasoning
WO2024238024A1
Object detection using visual language models via latent feature adaptation with synthetic data
WO2025106451A1