Urban management case image recognition method and system based on multi-modal large model
Patent Information
- Application Number
- CN202511401449.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-28
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2045-09-28
AI Technical Summary
[0005]为此,本发明实施例提供一种基于多模态大模型的城市管理案件图像识别方法及系统,以解决现有技术在易混淆的场景中,图片语义与文本语义难以对齐,识别精度不高的技术问题
[0041]本发明实施例通过采集城市管理图像及文本数据构建图像-文本对,利用多模态大模型为图像生成描述,并微调文本嵌入模型以提升图文语义对齐能力。本发明实施例可接受文本或图像查询通过混合检索策略结合稀疏检索与稠密检索优势,精准检索相关信息并最终利用多模态大模型进行深度推理。有效解决了图文语义未充分对齐和检索精度不足的问题显著提升了城市管理场景中复杂图像识别的准确性与可靠性。
Smart Images

Figure CN121074769B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image and text retrieval technology, specifically to a method and system for image recognition of urban management cases based on a multimodal large model. Background Technology
[0002] Currently, image recognition tasks in urban management largely rely on traditional image classification methods. These methods typically simplify the problem to a single-label classification task, failing to fully consider the possibility that a single image may contain multiple complex issues in real-world scenarios. For example, certain fine-grained problem types may appear visually similar, making accurate differentiation difficult through image classification alone, thus limiting recognition performance.
[0003] In recent years, with the development of multimodal large models, image and text retrieval technology has been gradually applied to cross-modal understanding tasks. Existing methods typically encode images and text into vector representations separately and match them by calculating the similarity between vectors. However, due to the inherent semantic gap between images and text, simple vector similarity matching often fails to achieve true semantic alignment. Especially when facing easily confused or diverse urban management scenarios, the performance of such methods is still not ideal.
[0004] Furthermore, existing multimodal retrieval systems often can only handle a single type of input, either text queries or image queries, lacking the ability to uniformly process multimodal inputs. This limits their flexibility and practicality in complex application scenarios. Summary of the Invention
[0005] To address this issue, this invention provides a method and system for image recognition of urban management cases based on a multimodal large model, in order to solve the technical problem that existing technologies struggle to align image semantics with text semantics in easily confused scenarios, resulting in low recognition accuracy.
[0006] To achieve the above objectives, the embodiments of the present invention provide the following technical solutions:
[0007] According to a first aspect of the present invention, a method for image recognition of urban management cases based on a multimodal large model is provided, the method comprising:
[0008] Image data of urban management cases were obtained through urban cameras and manual collection methods, and text descriptions of all problem types were compiled.
[0009] Attach a corresponding question type text description to each image to construct an image-text pair dataset;
[0010] A multimodal large model is used to generate text descriptions for the images in the image-text pair dataset, so that each image corresponds to both manually annotated text and model-generated text.
[0011] Using the manually annotated text and the model-generated text, combined with keyword information extracted from the manually annotated text, a text embedding model is fine-tuned to better align the semantics of the model-generated text with the manually annotated text.
[0012] Receive query information input by the user and detect the type of query information;
[0013] If the query information is text, the fine-tuned text embedding model is used to vectorize the query text, and a similarity search is performed in the image description vector database to obtain a preset number of the most relevant image descriptions and their associated images. The query text, the retrieved image descriptions and images are then input into the multimodal large model to output the final recognition result.
[0014] If the query information is an image, the query image is first input into the multimodal large model to generate its text description. Then, a hybrid retrieval strategy is used to process the generated text description to retrieve relevant text descriptions. The retrieved text, the description of the query image, and the query image itself are input into the multimodal large model to output the final recognition result.
[0015] Furthermore, using the manually annotated text and the model-generated text, combined with keyword information extracted from the manually annotated text, a text embedding model is fine-tuned to better align the semantics of the model-generated text with the manually annotated text, including:
[0016] A triplet loss function is constructed using manually annotated text as anchor points, model-generated text corresponding to the same image as positive samples, and text descriptions corresponding to other images in the same batch as negative samples.
[0017] Furthermore, when fine-tuning the text embedding model, a keyword supervision loss is introduced, which is constructed by calculating the distance between the embedding representation of manually annotated text and the embedding representation of its keywords.
[0018] Furthermore, the total loss function of the fine-tuned text embedding model is a weighted sum of the triplet loss and the keyword supervision loss.
[0019] Furthermore, before receiving user-input query information, a vector database construction phase is also included:
[0020] The fine-tuned text embedding model was used to vectorize the text descriptions of all question types and store them in the first vector database.
[0021] All images are Base64 encoded, and text descriptions for each image are generated using a multimodal large model.
[0022] The fine-tuned text embedding model is used to vectorize the text descriptions of all images and store them in a second vector database.
[0023] The Base64 encoded image data is stored in the storage database and associated with the vector description in the second vector database using a unique identifier.
[0024] Furthermore, when the query information is text, vector similarity retrieval is used when searching in the second vector database of image descriptions.
[0025] Furthermore, when the query information is an image, the hybrid retrieval strategy employed specifically includes:
[0026] Sparse retrieval is performed on the text descriptions generated from the query images to obtain the Top-K relevant text documents;
[0027] Extract keywords from the top-K relevant text documents and expand the query to obtain an expanded term set;
[0028] The original text description is concatenated with the expanded vocabulary to form a new query text.
[0029] The new query text is used to perform a dense search in Vector Database 1 to retrieve the most relevant question type text descriptions.
[0030] Furthermore, the multimodal large model is a pre-trained deep learning model capable of simultaneously processing and understanding image and text information.
[0031] According to a second aspect of the present invention, an image recognition system for urban management cases based on a multimodal large model is provided, the system comprising:
[0032] The data acquisition module is used to acquire image data of urban management cases through urban cameras and manual collection methods, and to organize text descriptions of all problem types;
[0033] The image-text pair dataset building module is used to attach a corresponding question type text description to each image to build the image-text pair dataset;
[0034] The text description generation module is used to generate text descriptions for the images in the image-text pair dataset using a multimodal large model, so that each image corresponds to both manually annotated text and model-generated text.
[0035] The model fine-tuning module is used to fine-tune a text embedding model by using the manually annotated text and the model-generated text, combined with keyword information extracted from the manually annotated text, so that the semantics of the model-generated text and the manually annotated text are better aligned.
[0036] The query information receiving module is used to receive query information input by the user and detect the type of query information;
[0037] The retrieval and identification module is used to perform the following steps:
[0038] If the query information is text, the fine-tuned text embedding model is used to vectorize the query text, and a similarity search is performed in the image description vector database to obtain a preset number of the most relevant image descriptions and their associated images. The query text, the retrieved image descriptions and images are then input into the multimodal large model to output the final recognition result.
[0039] If the query information is an image, the query image is first input into the multimodal large model to generate its text description. Then, a hybrid retrieval strategy is used to process the generated text description to retrieve relevant text descriptions. The retrieved text, the description of the query image, and the query image itself are input into the multimodal large model to output the final recognition result.
[0040] The embodiments of the present invention have the following advantages:
[0041] This invention constructs image-text pairs by collecting urban management images and text data, generates image descriptions using a multimodal large model, and fine-tunes the text embedding model to improve image-text semantic alignment. This invention accepts text or image queries, combines the advantages of sparse and dense retrieval using a hybrid retrieval strategy, accurately retrieves relevant information, and ultimately performs deep inference using the multimodal large model. It effectively solves the problems of insufficient image-text semantic alignment and inadequate retrieval accuracy, significantly improving the accuracy and reliability of complex image recognition in urban management scenarios. Attached Figure Description
[0042] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely exemplary, and those skilled in the art can derive other embodiments based on the provided drawings without creative effort.
[0043] The structures, proportions, sizes, etc. illustrated in this specification are only for the purpose of assisting those skilled in the art in understanding and reading the content disclosed herein, and are not intended to limit the conditions under which the present invention can be implemented. Therefore, they have no substantial technical significance. Any modifications to the structure, changes in the proportions, or adjustments to the size, without affecting the effects and objectives that the present invention can produce, should still fall within the scope of the technical content disclosed in the present invention.
[0044] Figure 1This is a schematic diagram of the logical structure of an urban management case image recognition system based on a multimodal large model, provided in an embodiment of the present invention.
[0045] Figure 2 A flowchart illustrating an image recognition method for urban management cases based on a multimodal large model, provided in an embodiment of the present invention;
[0046] Figure 3 This is a schematic diagram of the text recognition process in an image recognition method for urban management cases based on a multimodal large model, provided by an embodiment of the present invention.
[0047] Figure 4 This is a schematic diagram of the image class recognition process in an image recognition method for urban management cases based on a multimodal large model, provided in an embodiment of the present invention. Detailed Implementation
[0048] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0049] Currently, when image recognition issues are involved in urban management, traditional image classification methods are commonly used, and there is no evidence of modeling these issues as image-text retrieval problems. This approach is less effective in certain scenarios, such as when an image may contain multiple types of problems, or when some sub-problems are quite similar in the image.
[0050] Current image-text retrieval methods based on multimodal large models mainly vectorize text images through multimodal large models and then simply perform similarity matching. This approach makes it difficult to truly align image semantics with text semantics, resulting in poor performance in some easily confused scenarios.
[0051] To address the technical challenges of aligning image and text semantics in easily confused scenarios, resulting in low recognition accuracy, and to solve image recognition problems in urban management cases, such as damaged roads and manhole covers, damaged green facilities, and disorderly parking, etc.
[0052] refer to Figure 1 This invention discloses an image recognition system for urban management cases based on a multimodal large model. The system includes: a data acquisition module 1; an image-text pair dataset construction module 2; a text description generation module 3; a model fine-tuning module 4; a query information receiving module 5; and a retrieval and recognition module 6.
[0053] Corresponding to the aforementioned image recognition system for urban management cases based on a multimodal large model, this invention also discloses an image recognition method for urban management cases based on a multimodal large model. The following details the image recognition method for urban management cases based on a multimodal large model disclosed in this invention, in conjunction with the aforementioned image recognition system for urban management cases based on a multimodal large model.
[0054] refer to Figure 2 This invention discloses an image recognition method for urban management cases based on a multimodal large model. The method acquires image data of urban management cases through urban cameras and manual collection, and compiles text descriptions for all problem types. A corresponding text description for each problem type is attached to each image, constructing an image-text pair dataset. A multimodal large model is used to generate text descriptions for the images in the image-text pair dataset, ensuring that each image corresponds to both manually annotated text and model-generated text. Using the manually annotated text and model-generated text, combined with keyword information extracted from the manually annotated text, a text embedding model is fine-tuned to better align the semantics of the model-generated text with the manually annotated text.
[0055] 1) Collect image data through city cameras and manual methods. Compile text descriptions (text data) for all question types.
[0056] 2) Label the data, that is, attach a corresponding question type description to an image (there may be multiple questions), that is, construct a set of image-text pairs.
[0057] 3) Use a multimodal large model to generate text descriptions for images in the image-text pair data. In this way, one image corresponds to two text data: one is the text data labeled in step 2) (manually labeled), and the other is the text description generated by the multimodal large model.
[0058] 4) Fine-tune the text embedding model using the two text data points of the image and the keyword information of the labeled text data, so that the semantics of the text description generated by the multimodal large model can be better aligned with the original text annotations of the image. Specifically:
[0059] ① Data preparation: image-text pair data, image descriptions generated by multimodal large model, and keywords extracted from labeled text data by keyword extraction tools;
[0060] ② Constructing triplet: Anchor: labeled text data, positive sample: generated text description, negative sample: text description of other images in the batch, construct triplet loss;
[0061] ③ Integrate keyword information: Introduce keyword supervision loss, calculate the distance between the embedding of the labeled text data and its keyword embedding and minimize it;
[0062] ④ The weighted sum of triplet loss and keyword supervision loss is used as the total loss function to fine-tune the text embedding model.
[0063] 5) Use the fine-tuned embedding model to vectorize the text data and store it in the first vector database.
[0064] 6) Base64 encode the images, and then use a multimodal large model to semantically describe the images to form image descriptions. The image descriptions are vectorized and stored in a second vector database using a fine-tuned embedding model. The base64-encoded images are stored in another storage database (3), and images and their descriptions are associated using unique identifiers.
[0065] refer to Figure 3 and Figure 4 It receives query information input by the user and detects the type of query information.
[0066] refer to Figure 3 If the query information is text, the fine-tuned text embedding model is used to vectorize the query text, and a similarity search is performed in the image description vector database to obtain a preset number of image descriptions with the highest relevance and their associated images. The query text, the retrieved image descriptions and images are then input into the multimodal large model to output the final recognition result.
[0067] refer to Figure 4 If the query information is an image, the query image is first input into the multimodal large model to generate its text description. Then, a hybrid retrieval strategy is used to process the generated text description to retrieve the relevant text description. The retrieved text, the description of the query image, and the query image itself are input into the multimodal large model to output the final recognition result.
[0068] 7) Accept user input, which can be text (e.g., querying whether a certain type of question exists) or images (e.g., querying what questions are contained in an image). If querying whether a certain type of question exists, first, the question is vectorized using a fine-tuned embedding model. Then, relevant content is retrieved from the image description vector database, returning the most relevant image descriptions and corresponding images. Finally, the question, image descriptions, and images are fed into the multimodal large-scale model to provide the final result. If querying what questions are contained in an image, first, the image is converted to base64 encoding. Then, the multimodal large-scale model provides an image description. A hybrid retrieval strategy combining sparse retrieval and vector similarity retrieval is applied to the obtained image description. The retrieved text (i.e., the textual description of the question type), image description, sparse retrieval keywords and extended terms, and the image are fed into the multimodal large-scale model to provide the final result.
[0069] 8) The hybrid retrieval strategy using sparse retrieval combined with vector similarity retrieval described in the previous step is not a simple merging of the two retrieval results. Specifically:
[0070] ① Use the generated text description to quickly perform sparse retrieval to obtain the Top-K relevant documents;
[0071] ② Extract keywords from these Top-K documents and calculate expanded terms using a query expansion model (such as RM3);
[0072] ③ Patch the keywords obtained in the previous step to the end of the text description to form a new query text, which is then sent to the dense retrieval model for retrieval.
[0073] The embodiments of the present invention have the following beneficial effects:
[0074] This invention can receive both text and image input, and can handle a wider range of problems compared to methods that can only receive a single type of input.
[0075] This invention utilizes multimodal data to guide the fine-tuning of the text embedding model, so that the semantics of the text description generated by the multimodal large model can be better aligned with the original text annotation of the image;
[0076] This invention uses the results obtained from sparse retrieval to empower dense retrieval instead of simply merging the results of the two retrievals, thereby alleviating the word mismatch problem, enhancing the dense model's ability to capture key information in the query, and improving the accuracy of the retrieval results;
[0077] This invention fully utilizes the powerful image-text understanding capabilities of multimodal large models to generate text descriptions for images and employs an innovative hybrid retrieval strategy during the retrieval stage. Compared to directly using multimodal large models to convert images into vectors and then matching them with text vectors, this approach better aligns the semantics of images and text, resulting in better recognition performance.
[0078] This invention performs more refined context engineering on large models, enabling them to generate better responses.
[0079] Although the present invention has been described in detail above with general descriptions and specific embodiments, modifications or improvements can be made to it, which will be obvious to those skilled in the art. Therefore, all such modifications or improvements made without departing from the spirit of the present invention fall within the scope of protection claimed by the present invention.
Claims
1. A method for image recognition of urban management cases based on a multimodal large model, characterized in that, The method includes: Image data of urban management cases were obtained through urban cameras and manual collection methods, and text descriptions of all problem types were compiled. Attach a corresponding question type text description to each image to construct an image-text pair dataset; A multimodal large model is used to generate text descriptions for the images in the image-text pair dataset, so that each image corresponds to both manually annotated text and model-generated text. Using the manually annotated text and the model-generated text, combined with keyword information extracted from the manually annotated text, a text embedding model is fine-tuned to better align the semantics of the model-generated text with the manually annotated text. Receive query information input by the user and detect the type of query information; If the query information is text, the fine-tuned text embedding model is used to vectorize the query text, and a similarity search is performed in the image description vector database to obtain a preset number of the most relevant image descriptions and their associated images. The query text, the retrieved image descriptions and images are then input into the multimodal large model to output the final recognition result. If the query information is an image, the query image is first input into the multimodal large model to generate its text description. Then, a hybrid retrieval strategy is used to process the generated text description to retrieve the relevant text description. The retrieved text, the description of the query image, and the query image itself are input into the multimodal large model to output the final recognition result. Among them, the relevant text descriptions retrieved include keywords and their extended words extracted from documents obtained by sparse retrieval and question type descriptions obtained by dense retrieval in the first vector database; Using the manually annotated text and model-generated text, and incorporating keyword information extracted from the manually annotated text, a text embedding model is fine-tuned to better align the semantics of the model-generated text with the manually annotated text. This includes: A triplet loss function is constructed using manually annotated text as anchor points, model-generated text corresponding to the same image as positive samples, and text descriptions corresponding to other images in the same batch as negative samples. When fine-tuning the text embedding model, a keyword supervision loss was also introduced, which was constructed by calculating the distance between the embedding representation of manually annotated text and the embedding representation of its keywords. The total loss function of the fine-tuned text embedding model is a weighted sum of the triplet loss and the keyword supervision loss; When the query information is an image, the hybrid retrieval strategy employed specifically includes: Sparse retrieval is performed on the text descriptions generated from the query images to obtain the Top-K relevant text documents; Extract keywords from the top-K relevant text documents and expand the query to obtain an expanded term set; The original text description is concatenated with the expanded vocabulary to form a new query text. The new query text is used to perform a dense search in the first vector database to retrieve the most relevant question type text descriptions.
2. The method for image recognition of urban management cases based on a multimodal large model as described in claim 1, characterized in that, Before receiving user-input query information, there is also a vector database construction phase: The fine-tuned text embedding model was used to vectorize the text descriptions of all question types and store them in the first vector database. All images are Base64 encoded, and text descriptions for each image are generated using a multimodal large model. The fine-tuned text embedding model is used to vectorize the text descriptions of all images and store them in a second vector database. The Base64 encoded image data is stored in the storage database and associated with the vector description in the second vector database using a unique identifier.
3. The method for image recognition of urban management cases based on a multimodal large model as described in claim 2, characterized in that, When the query information is text, vector similarity retrieval is used when searching in the second vector database of image descriptions.
4. The method for image recognition of urban management cases based on a multimodal large model as described in claim 1, characterized in that, The multimodal large model is a pre-trained deep learning model capable of simultaneously processing and understanding image and text information.
5. A system for recognizing urban management cases based on a multimodal large model, characterized in that, The system includes: The data acquisition module is used to acquire image data of urban management cases through urban cameras and manual collection methods, and to organize text descriptions of all problem types; The image-text pair dataset building module is used to attach a corresponding question type text description to each image to build the image-text pair dataset; The text description generation module is used to generate text descriptions for the images in the image-text pair dataset using a multimodal large model, so that each image corresponds to both manually annotated text and model-generated text. The model fine-tuning module is used to fine-tune a text embedding model by using the manually annotated text and the model-generated text, combined with keyword information extracted from the manually annotated text, so that the semantics of the model-generated text and the manually annotated text are better aligned. The query information receiving module is used to receive query information input by the user and detect the type of query information; The retrieval and identification module is used to perform the following steps: If the query information is text, the fine-tuned text embedding model is used to vectorize the query text, and a similarity search is performed in the image description vector database to obtain a preset number of the most relevant image descriptions and their associated images. The query text, the retrieved image descriptions and images are then input into the multimodal large model to output the final recognition result. If the query information is an image, the query image is first input into the multimodal large model to generate its text description. Then, a hybrid retrieval strategy is used to process the generated text description to retrieve the relevant text description. The retrieved text, the description of the query image, and the query image itself are input into the multimodal large model to output the final recognition result. Using the manually annotated text and model-generated text, and incorporating keyword information extracted from the manually annotated text, a text embedding model is fine-tuned to better align the semantics of the model-generated text with the manually annotated text. This includes: A triplet loss function is constructed using manually annotated text as anchor points, model-generated text corresponding to the same image as positive samples, and text descriptions corresponding to other images in the same batch as negative samples. When fine-tuning the text embedding model, a keyword supervision loss was also introduced, which was constructed by calculating the distance between the embedding representation of manually annotated text and the embedding representation of its keywords. The total loss function of the fine-tuned text embedding model is a weighted sum of the triplet loss and the keyword supervision loss; When the query information is an image, the hybrid retrieval strategy employed specifically includes: Sparse retrieval is performed on the text descriptions generated from the query images to obtain the Top-K relevant text documents; Extract keywords from the top-K relevant text documents and expand the query to obtain an expanded term set; The original text description is concatenated with the expanded vocabulary to form a new query text. The new query text is used to perform a dense search in the first vector database to retrieve the most relevant question type text descriptions.
Citation Information
Patent Citations
Scene text detection and recognition method based on cross-modal large language model
CN117851883A
Multi-modal natural language understanding and generating system and method
CN120670635A