Medium-high point image scene understanding method and system based on scene interpretation graph
Through the method based on scene interpretation diagram, using text prompt words and large language models to generate fine scene interpretation diagrams, the problems of large data requirements and unexplainable results in traditional methods are solved, and efficient and accurate understanding of mid- and high-point image scenes are achieved.
Patent Information
- Application Number
- CN202510811655.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-18
AI Technical Summary
Traditional medium and high point image scene understanding methods require a large amount of training data, which is costly and long iteration cycles, and the result interpretability is weak. The square box of the detection result ignores the target details, resulting in errors in the relationship calculation, affecting the accuracy of scene understanding.
By inputting text prompt words and scene images into the open set object detection model, obtaining the object detection box and label, using the segmentation model to obtain the mask, calculate the relative position and size relationship of the target, and using the large language model to generate a scene interpretation map and output text results.
It improves the generalization ability of scenario understanding, saves data training costs, enhances the accuracy and interpretability of descriptions, and ensures the accuracy and logic of scenario understanding results.
Smart Images

Figure CN120340031A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of graphic and text data processing, and specifically to a method and system for middle and high point image scene understanding based on a scene interpretation map. Background Art
[0002] With the development of information technology, especially the progress of computer vision, big data, artificial intelligence and Internet technology, the scene monitoring and scene understanding technology for middle and high points has also been significantly developed. The data volume collected by middle and high point cameras and various sensors has increased extremely rapidly, storing a large amount of characteristic data of middle and high points with industry attributes such as emergency, forestry and grassland, land and resources, water conservancy, environmental protection, and agriculture, providing data support for training to obtain scene understanding of middle and high point images.
[0003] Traditional methods for middle and high point image scene understanding use labeled scene data to classify scene categories, divide the scene into categories such as forest, farmland, factory, etc., obtain a scene classification model by using the method of training a neural network model, use this model to classify the images in actual applications, and finally obtain the category result of the scene; then, a scene understanding method based on a large number of training data target detection models uses a large amount of image annotation data to train and obtain a target detection model, and the model detects the input image and directly uses the detection result as the result of scene image understanding; furthermore, a scene understanding method based on a large number of text images for training data uses a large amount of image and text annotation information, obtains an image understanding model through training, and in the process of scene understanding, the model directly inputs the image and outputs the text explanation of the scene, and uses this text as the result of scene understanding; finally, a scene understanding method through self-supervised learning uses a large amount of unlabeled data, and through self-supervised learning technology, the image understanding model learns image features without manual annotation. Traditional methods for middle and high point image scene understanding need to collect a large amount of scene data, annotate the data, and train the model through the annotated data. In the inference process, by inputting the image, the detection target box information with fixed labels is obtained. It is also necessary to use a large amount of data to train and obtain a large vision model, and in the inference process, input the image to directly obtain the classification result of the image and the text understanding content of the image.
[0004] This requires a large amount of training data. After the model is trained, the application boundary of the model is determined accordingly. When the application scenario changes, the model's capabilities will weaken, and more training data needs to be introduced for training. In the scene understanding method based on large models, training a large model requires massive amounts of data and huge training resources. Therefore, the R & D cost of using this method is high, and the iteration cycle is long. When using a vision large model for scene understanding, the model directly gives text results, and the intermediate reasons and judgment bases cannot be obtained. When unreasonable results are generated, the reasons cannot be found. Therefore, using a vision large model for scene understanding has the disadvantages of weak interpretability and non-retraceability. In traditional scene interpretation graph analysis, the square boxes of the detection results are used. Since the shapes of the targets are irregular, in the process of using the square boxes to represent the targets for analysis, the details of the targets are ignored, resulting in incorrect calculation of the relationships between objects and affecting the final scene understanding results. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for understanding medium and high point images based on a scene interpretation graph, which improves the generalization ability of scene understanding, saves data training costs, improves the accuracy and detail of descriptions, and improves the interpretability of understanding results.
[0006] The purpose of the present invention can be achieved through the following technical solutions: A method for understanding medium and high point images based on a scene interpretation graph includes the following steps: Input text prompt words and a scene image into an open-set object detection model to obtain the object detection boxes and object labels required for detecting scene understanding; Input the object detection boxes and the scene image into a segmentation model to obtain the masks of the object detection boxes; Calculate the relative positions and size relationships of the objects through the labels and masks to obtain a scene interpretation graph; Analyze the scene interpretation graph through a large language model and output text results for image scene understanding.
[0007] In a further solution, the method of inputting text prompt words and a scene image into the open-set object detection model includes: Obtain a scene image; Preset a text description of the scene to be analyzed for the scene image; Input the text description into a large language model, and generate the text prompt words through the large language model.
[0008] In a further solution, the method of calculating the relative positions and size relationships of the objects through the labels and masks to obtain a scene interpretation graph includes: Obtain the central position, boundary contour shape, boundary length, and boundary pixel area of the object detection box mask; Calculate the relative position and size relationship between the target and other objects in the scene image based on the central position, boundary contour shape, boundary length, and boundary pixel area; Establish a scene interpretation map based on the relative position and size relationship between the target and other objects in the scene image and the labels to represent the relationships between objects.
[0009] In a further aspect, the method for analyzing the scene interpretation map by the large language model and outputting the text result for image scene understanding includes: Input the position relationship, inclusion relationship, relative area size relationship between the target and other objects in the scene image, and the target label into the large language model, and output the text result for image scene understanding through the large speech model.
[0010] In a further aspect, the method further includes: Input the questions related to the image scene understanding into the large language model; The large language model continues to reason based on the output text result for image scene understanding, obtains the understanding result of the related questions, and outputs feedback to the user.
[0011] In a further aspect, the relative position relationship includes the relative distance of the central position, the relative far and near distance of the boundary contour, and the intersection relationship of the boundary contours, and the relative size relationship includes the relative area relationship and the boundary contour length relationship.
[0012] In a further aspect, the text prompt words are different key prompt words proposed according to the application scene image and the problem to be analyzed.
[0013] The present invention also proposes a high and middle point image scene understanding system based on a scene interpretation map, including: A target detection box and target label acquisition module, configured to input text prompt words and a scene image into an open-set target detection model to obtain the target detection box and target label required for detecting scene understanding; A mask acquisition module for the target detection box, configured to input the target detection box and the scene image into a segmentation model to obtain the mask of the target detection box; A scene interpretation map acquisition module, configured to calculate the relative position and size relationship of the target through the label and the mask to obtain a scene interpretation map; An image scene understanding output module, configured to analyze the scene interpretation map through the large language model and output the text result for image scene understanding.
[0014] In a further aspect, the open-set target detection model is the Grounding DINO model or the YOLO model.
[0015] In a further solution, the segmentation model is a model of the Segment Anything Model series.
[0016] Advantages of the present invention: Through the description of the scenario requirements, the present invention automatically generates text prompts to be detected by using a large language model, and completes object detection by using a general open-set object detection model. This method can automatically complete the detection of key objects according to different scenarios and requirements, and has the advantages of strong adaptability and wide generalization.
[0017] By detecting the labels of the objects, using a segmentation model to complete object segmentation, and calculating the orientation relationship, inclusion relationship, proximity relationship, etc. of the objects by using a refined mask, a scene interpretation map is finally obtained. Compared with the square bounding boxes used in the scene interpretation map in the existing methods, it has a finer granularity, so the calculation results are also more accurate, making the final scene understanding result more accurate. This method calculates the scene interpretation map by using the mask of the object, and has the advantages of detailed description and strong accuracy; The present invention converts the scene interpretation map into text information, and completes the understanding and induction of the scene content through a large language model, and finally generates answers and disposal measures corresponding to the requirements. This method uses the refined scene interpretation map as the input of the large language model, and has the advantages of strong interpretability and strong generalization. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1 is a flowchart of a method for understanding a medium-high point image scene based on a scene interpretation map in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0021] As Figure 1 shown, a method for understanding a medium-high point image scene based on a scene interpretation map includes the following steps: S1: Input the text prompt and the scene image into the open-set object detection model to obtain the object detection bounding boxes and object labels required for understanding the detection scene. The open-set object detection model can be the open-set object detection model GroundingDINO or the YOLO series models, etc. The input text prompt can be the objects to be detected, such as farmland, forest, road, and / or pedestrian, etc. Input these text prompts into the open-set object detection model to obtain the object detection bounding boxes, that is, the bounding boxes corresponding to the input text prompts. These bounding boxes are the elemental objects required for understanding the detection scene, and obtain the labels of these objects to distinguish the categories of the objects.
[0022] S2: Input the object detection bounding boxes and the scene image into the segmentation model to obtain the masks of the object detection bounding boxes. The segmentation model can be models such as the Segment Anything Model series. Use this segmentation model to separate the detected object detection bounding boxes in step S1 above to obtain the masks corresponding to each bounding box, and obtain the outer contours at the target pixel level.
[0023] S3: Calculate the relative positions and size relationships of the objects through the labels and masks to obtain the scene interpretation map. S4: Analyze the scene interpretation map through the large language model and output the text results for understanding the image scene.
[0024] According to the above embodiments, in some embodiments, step S1 further includes: Obtain the scene image. The method of obtaining the image of the scene can be to obtain it from historical data or to take a current photo.
[0025] Preset the text description of the scene to be analyzed for the scene image. For example: I am now dealing with a fire scene, and the preset text description can be: I now need to detect the smoke and fire in the image. This image is an urban scene, and I need to obtain the targets of the smoke and fire and other relevant targets in this scene.
[0026] Input the text description into the large language model, and generate the text prompt through the large language model. The large language model will return text prompts related to smoke and fire, buildings, pedestrians, roads, chimneys, factories, etc.
[0027] In some embodiments, step S3 includes: Obtain the central position, boundary contour shape, boundary length, and boundary pixel area of the object detection bounding box mask; Calculate the relative positions and size relationships between the object and other objects in the scene image through the central position, boundary contour shape, boundary length, and boundary pixel area; A scene interpretation graph is established based on the relative position, size relationship, and labels between the target and other objects in the scene image to represent the connections between objects. For example, the obtained scene interpretation graph may show that two targets, a farmland and a fire, do not intersect, are relatively far apart in the graph, the area of the fire is small, and the area of the farmland is large. Similarly, the relationship between each target and other targets can be obtained. The angular azimuth position relationship (up, down, left, right, and angular relationship) and inclusion relationship (intersection, inclusion, and separation relationship) between the target and other scene objects are calculated through the mask boundary contour, the distance relationship between the target and other scene objects is calculated through the mask center position, and the relative area size relationship between the target and other scene objects is calculated through the pixel area occupied by the mask boundary. Based on the relationship between the target and other scene objects and the detection labels, a high-precision scene interpretation graph can be established to represent the connections between objects. A detailed scene interpretation graph can help the large language model understand the scene information more accurately.
[0028] In some embodiments, step S4 includes: Input the position relationship, inclusion relationship, relative area size relationship, and target label between the target and other objects in the scene image into the large language model, and the large language model outputs the text result of the image scene understanding. In some embodiments, the method further includes step S5: Input the questions related to the image scene understanding into the large language model; The large language model continues to reason based on the output text result of the image scene understanding, obtains the understanding result of the relevant questions, and outputs feedback to the user.
[0029] For example, the results in the previously obtained scene interpretation graph are as follows: There are 2 farmlands and 2 fires in the scene. Both fires are within the range of the first farmland, and the areas of both fires are relatively small, etc. There are many relationships. The user question that the large language model needs to solve for this scene is: Is there a fire happening now? What should I do? Input the results of the scene interpretation graph (in text form) and the question to be solved (in text form) into the large language model, and the following text feedback will be given: This scene is a farmland scene where a fire is occurring. There are two fire points in total, located in the first farmland (and the coordinates xy are marked). Please contact the safety officer in time to extinguish the fire.
[0030] For the scene interpretation graph obtained in step S3, input the relationships of this graph in text form into the large language model. Utilizing the summarization ability of the large language model, the overall understanding and description results of this scene can be output. At the same time, if there are other related questions about this scene image, during the reasoning process, input the question and the results of the scene interpretation graph into the large language model together, and the understanding results of the corresponding questions can be obtained.
[0031] In some embodiments, the relative position relationship includes the relative distance of the central positions, the relative far - near distances of the boundary contours, and the intersection relationship of the boundary contours, and the relative size relationship includes the relative area relationship and the boundary contour length relationship.
[0032] In some embodiments, the text prompt words are different key prompt words proposed according to the application - scenario image and the problem to be analyzed.
[0033] According to the above - mentioned method, an embodiment can be obtained. In the embodiment, a high - point image scene understanding system based on a scene interpretation graph includes: A target detection box and target label acquisition module, configured to input text prompt words and a scene image into an open - set target detection model, and acquire the target detection box and target labels required for detecting scene understanding; A mask acquisition module for the target detection box, configured to input the target detection box and the scene image into a segmentation model to acquire the mask of the target detection box; A scene interpretation graph acquisition module, configured to calculate the relative position and size relationship of the target through the label and the mask, and obtain the scene interpretation graph; An image scene understanding output module, configured to analyze the scene interpretation graph through a large - language model and output a text result for image scene understanding.
[0034] Compared with traditional scene understanding methods, the method of the present invention has the following improvements: 1. Using the above - mentioned technical solution can make the method of the present invention have stronger application generalization. In different application scenarios, different keywords can be extracted according to the characteristics of the application - scenario image and the problem to be analyzed. The method of the present invention will generate different scene interpretation graphs according to different keywords, and thus different scene understanding results will be obtained. Compared with existing methods, when changing the application scenario, good application results can be obtained without more training data.
[0035] 2. Using the above - mentioned technical solution can make the method of the present invention have stronger logic. In the process of scene understanding, the method of the present invention will obtain the result of scene understanding according to the calculated scene interpretation graph. The basis is clear, and through precise relational numerical calculation, the result of scene understanding has stronger logic and traceability.
[0036] 3. Using the above - mentioned technical solution can make the method of the present invention have a more fine - grained process and more accurate scene understanding. The method of the present invention uses a more precise target mask to calculate the scene interpretation graph. Compared with the square bounding box used in the scene interpretation graph of existing methods, it has a finer granularity, so the calculated result is also more accurate, making the final scene understanding result more accurate.
[0037] It should be noted that the terms "first", "second", etc. in this application are used to distinguish similar objects and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances for the embodiments of the present application described herein.
[0038] In the description of this specification, the description with reference to terms such as "one embodiment", "example", "specific example", etc. means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0039] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art of this industry should understand that the present invention is not limited by the above embodiments, and what is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed.
Claims
1. A method for understanding the image scene of medium and high points based on a scene interpretation graph, characterized in that, It includes the following steps: Input text prompts and a scene image into an open-set object detection model to obtain the object detection bounding boxes and object labels required for understanding the detection scene; Input the object detection bounding boxes and the scene image into a segmentation model to obtain the masks of the object detection bounding boxes; Calculate the relative positions and size relationships of the objects through the labels and masks to obtain a scene interpretation map; Analyze the scene interpretation map through a large language model and output a text result for understanding the image scene; The method of inputting text prompts and a scene image into the open-set object detection model includes: Obtain a scene image; Preset a text description of the scene to be analyzed for the scene image; Input the text description into a large language model, and generate the text prompts through the large language model.
2. The method for mid-high point image scene understanding based on a scene interpretation graph according to claim 1, wherein The method of calculating the relative positions and size relationships of the objects through the labels and masks to obtain a scene interpretation map includes: Obtain the central position, boundary contour shape, boundary length, and boundary pixel area of the object detection bounding box mask; Calculate the relative positions and size relationships between the object and other objects in the scene image through the central position, boundary contour shape, boundary length, and boundary pixel area; Establish a scene interpretation map through the relative positions and size relationships between the object and other objects in the scene image and the labels to represent the relationships between the objects.
3. A method for mid-high point image scene understanding based on a scene interpretation map according to claim 2, characterized in that The method of analyzing the scene interpretation map through a large language model and outputting a text result for understanding the image scene includes: Input the position relationships, inclusion relationships, relative area size relationships, and object labels between the object and other objects in the scene image into a large language model, and output the text result for understanding the image scene through the large language model.
4. The method for mid-high point image scene understanding based on a scene interpretation map according to claim 3, characterized in that, The method further includes: Input the questions related to understanding the image scene into the large language model; The large language model continues to reason based on the output text result for understanding the image scene, obtains the understanding results of the relevant questions, and outputs feedback to the user.
5. The method for understanding the medium and high point image scene based on the scene interpretation graph according to claim 2, characterized in that, The relative position relationships include the relative distance of the central positions, the relative far and near distances of the boundary contours, and the intersection relationships of the boundary contours. The relative size relationships include the relative area relationship and the boundary contour length relationship.
6. The method for mid-high point image scene understanding based on a scene interpretation graph according to claim 1, characterized in that, The text prompts are different key prompts proposed according to the application scene image and the questions to be analyzed.
7. A mid-high point image scene understanding system based on a scene interpretation graph, characterized in that, It includes: An object detection bounding box and object label acquisition module, which is used to input text prompts and a scene image into an open-set object detection model to obtain the object detection bounding boxes and object labels required for understanding the detection scene; An object detection bounding box mask acquisition module, which is used to input the object detection bounding boxes and the scene image into a segmentation model to obtain the masks of the object detection bounding boxes; A scene interpretation map acquisition module, which is used to calculate the relative positions and size relationships of the objects through the labels and masks to obtain a scene interpretation map; An image scene understanding output module, which is used to analyze the scene interpretation map through a large language model and output a text result for understanding the image scene; The method of inputting text prompts and a scene image into the open-set object detection model includes: Obtain a scene image; Preset a text description of the scene to be analyzed for the scene image; Input the text description into a large language model, and generate the text prompts through the large language model.
8. An image scene understanding system for medium and high points based on a scene interpretation map according to claim 7, characterized in that, The open-set object detection model is the Grounding DINO model or the YOLO model.
9. The scene understanding system for medium and high point images based on a scene interpretation graph according to claim 7, characterized in that, The segmentation model is the Segment Anything Model series of models.
Citation Information
Patent Citations
Scene description information determination method and device based on scene feature extraction
CN113269088A
Content understanding method and device for video static scene, equipment and medium
CN117975336A
Instance-level scene recognition using visual language models
CN118587623A
Multi-modal large model scene understanding method based on scene graph enhancement
CN119418339A
Image tag understanding method and device, electronic equipment and storage medium
CN119445203A
Cited By
Steel ladle hoisting safety monitoring and early warning method, system, equipment and medium
CN121191081A