Image-text retrieval graph network method based on model multiplexing
By constructing a graphic and text search network method, using pre-trained models and graph network modeling, the representation conflict problem in the graphic and text search tasks in professional fields is solved, and efficient and accurate image and text fine-grained correlation is achieved, which improves the search effect.
Patent Information
- Application Number
- CN202510520030.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art has characterization conflicts in the graphic and text retrieval tasks in the professional field, and it is difficult to adapt to the fine-grained correlation between images and text in the professional field, resulting in poor retrieval results.
By constructing a graph-text search network method, image meta features are extracted using pre-trained models, combined with graph network modeling and text adaptive image feature extraction, fine-grained image features are generated, and adjacent images provide semantic information to supplement the search context.
It realizes efficient and accurate graphic and text retrieval in the professional field, improves the fine-grained feature representation ability of retrieval, and improves the accuracy and efficiency of retrieval.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] The present invention relates to a graph network method for image-text retrieval based on model reuse, which is suitable for intelligent retrieval and model reuse of massive heterogeneous data in industrial scenarios, especially for scenarios such as image retrieval, e-commerce recommendation and multimedia content analysis, and belongs to the field of computer vision and cross-modal retrieval technology. Background Art
[0002] As a core technological breakthrough in the field of intelligent recommendation, the image-text cross-modal retrieval system demonstrates strong application value in scenarios such as precise e-commerce recommendation and social media content understanding by deeply integrating visual and semantic features. The system can efficiently handle compound query requirements of "image search + text modification". For example, based on clothing pictures uploaded by users and combined with text modification conditions such as "pure cotton material" and "spring new style", it can accurately locate target products from a large number of candidate products, and dynamically adapt to the retrieval needs of new categories and popular elements. Its core technology lies in building a unified multimodal representation space, realizing fine-grained alignment of local image features and text keywords through a hierarchical attention mechanism, and using an efficient approximate nearest neighbor search algorithm to screen out the retrieval results that best meet user intentions from tens of millions of image libraries while ensuring millisecond response speeds, greatly improving the conversion rate of e-commerce platforms and user shopping experience.
[0003] The current mainstream image-text cross-modal retrieval technology is mainly based on large-scale pre-trained multimodal models. These models can learn general visual-language aligned representations by pre-training on massive image-text datasets on the Internet. However, this "big data driven" training paradigm has obvious domain adaptability problems: the image features extracted by the general pre-training model are usually fixed, and it is difficult to focus on the image details based on the text information, resulting in the problem of "representation conflict" when directly reusing the pre-training model, which makes it difficult to adapt to the image-text retrieval task. Optimizing the model according to the "pre-training-fine-tuning" paradigm will not only destroy the original cross-modal association ability of the pre-training model encoder, but also may introduce noise or fall into local optimality due to adjusting the initial weights of the model, so that the final constructed image-text retrieval graph network loses the robustness of the general model in professional field retrieval tasks, and it is difficult to form accurate feature representations in professional field image-text retrieval tasks. In industrial image and text retrieval scenarios, it is necessary to use text information to focus on the local features of the image. Therefore, it is necessary to effectively model the fine-grained topological association between images and conditional text in professional fields to avoid the problem of missing semantic details when the model organizes professional domain knowledge of image and text retrieval tasks (such as the correspondence between pathological description nodes in medical imaging reports and visual feature nodes of lesion areas). Summary of the invention
[0004] Objective of the Invention: Since the previous methods lack the ability to structurally model the graphic and text retrieval tasks in professional fields and are difficult to cope with the "representation conflict" dilemma faced by general models in professional graphic and text retrieval tasks, the invention provides a graphic and text retrieval graph network method based on model reuse, which can achieve "efficient reuse of pre-trained models in specific retrieval tasks". Specifically, first, collect the data set of the proprietary field according to the user's needs, and reuse the pre-trained model to extract the meta-features of the data set, avoiding the high computational cost brought by training from scratch or fine-tuning all parameters. Then, use the graph network to model the topological structure of the image and automatically focus on the key regions of the image related to the text description, generate fine-grained image features, and complete the supplementary context information of the retrieval by means of the semantic information provided by similar images, enhancing the fine-grained features on the proprietary field data set, thereby assisting in the completion of the retrieval task. In this process, there is no need to fine-tune the pre-trained model, and only a graph neural network with few parameters needs to be introduced to efficiently complete the downstream task.
[0005] Technical Solution: A graphic and text retrieval graph network method based on model reuse includes three steps: image and text data acquisition, construction of a text-adaptive image feature extraction model, and retrieval of neighboring images; First, collect image data with different distributions and categories from the image and text data sets in the public data source, and each image is equipped with semantically related titles, labels or descriptive texts; Build a multi-level data preprocessing pipeline to denoise and filter the data in the original image and text data set; Then, build a text-adaptive image feature extraction model. The purpose of this step is to selectively extract the relevant features of the image according to the input text conditions. The specific process of building the text-adaptive image feature extraction model is to first use the open-source pre-trained model to extract the meta-features of the image and text, construct the initial node representation, design the topological structure of a single image, where the segmented regions of the image respectively form homogeneous subgraphs, and the cross-modal edges are dynamically generated through the graph attention mechanism; Generate condition text-adaptive image features; Finally, the specific process of retrieving neighboring images is to use the image nodes and text nodes to respectively form homogeneous subgraphs, and based on the graph sampling and clustering algorithm, obtain similar images of the query image to supplement the context information, and generate the target image features to complete the retrieval.
[0006] The specific steps of image and text data set acquisition are as follows: Step 100, construct an image and text data acquisition system, obtain the original image information from the public data source, and record the meta-information of the image at the same time.
[0007] Step 101, obtain the marked text of the image at the same network source address as the original image information. Based on the semantic association rule, synchronously obtain multi-dimensional marked texts strongly related to the image from the same network source address, including but not limited to user annotations, titles, descriptive texts, and scene labels.
[0008] Step 102: Adopt a three-level data cleaning and filtering mechanism, namely automated data extraction, pre-trained model for filtering semantic relationships, and expert review. Specifically, for the mechanism of filtering semantic relationships by the pre-trained model, deploy a pre-trained multi-modal model based on the Transformer architecture, achieve automated semantic matching by calculating the cross-modal similarity score of image-text pairs, and set a threshold to complete the preliminary filtering of low-quality samples. The expert review process is for domain experts to conduct sampling detection on the results to ensure that the obtained graphic and text semantic alignment is achieved and there are no problems such as low pixels and low quality in the image data.
[0009] Step 103: Use a hierarchical clustering algorithm to perform multi-granularity semantic clustering on the image dataset, and optimize sampling based on the intra-class diversity and inter-class discrimination metrics to ensure the balance and representativeness of the final dataset in both the visual semantics and text description dimensions.
[0010] After the above steps 100 - 103, an effective training image-text dataset is obtained.
[0011] The specific process steps for building a text-adaptive image feature extraction model are as follows: Step 200: Directly use the publicly available industrially pre-trained model to extract the patch features and text features of the images in the training image-text dataset obtained in the previous step. The patch features are non-overlapping image patches divided into 16 * 16. This step aims to provide more fine-grained image meta-features for this text-adaptive image feature extraction model. Step 201: Construct an image patch topology graph based on spatial relationships, where each image patch serves as a graph node, and edge relationships are established based on the 4-neighborhood connection strategy to form a regular grid graph structure. Step 202: Introduce the text features as global context nodes into the graph structure, and establish cross-modal associations with all image patch nodes through full connection to construct a heterogeneous base graph.
[0012] Step 203: Adopt a multi-head graph attention mechanism, set the dual-constraint contrastive learning loss function of single-modal and dual-modal as the training objective, and complete the training with the intra-image single-modal contrast loss and the image-text dual-modal contrast loss as the objectives. Based on the heterogeneous base graph generated in step 202, update the parameters of the multi-head graph attention module to generate text-conditioned adaptive image feature representations.
[0013] Based on the image features extracted in the previous step, the steps for retrieving neighboring images are as follows: Step 300: Construct an image semantic relationship graph based on feature similarity, and use cosine similarity to measure the semantic association strength of the fused features. Step 301: Adopt an importance-based random walk sampling algorithm to obtain the k-nearest neighbors of the query image starting from the query node.
[0014] Step 302: Design a multi-channel feature aggregation module: Proximity feature attention weighted aggregation and max pooling are used to capture significant feature aggregation modules. Finally, the features extracted by weighted aggregation and the features captured by max pooling are dynamically fused through a gating mechanism.
[0015] Step 303: Construct a dataset required for sample evaluation, including: 1) a candidate image library; 2) a reference image set; 3) multi-modal query conditions.
[0016] Step 304: Use the approximate nearest neighbor search algorithm to quickly match similar images in the feature library.
[0017] In the above step 203, the text-adaptive image feature model, and in steps 302 and 303, the graph sampling and aggregation module both adopt a hierarchical feature fusion architecture: First, through a heterogeneous multi-layer graph attention network, cross-modal interaction is performed between the original image patch features and text features to generate text-adaptive image features. Subsequently, a graph sampling and aggregation algorithm is designed to complete the context supplement of neighboring samples for the query. In particular, as mentioned in step 203, a contrastive learning objective function with double constraints is designed, including an intra-modal contrast loss and a cross-modal contrast loss, to respectively maintain the image structure characteristics and text semantic consistency, and finally, joint embedding features are output.
[0018] A computer device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the above computer program, it implements the steps of the above-described graph network method for image-text retrieval based on model reuse.
[0019] A computer-readable storage medium stores a computer program for executing the above-described graph network method for image-text retrieval based on model reuse. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 It is a flowchart of image-text dataset collection according to an embodiment of the present invention; Figure 2 It is a flowchart for building a text-adaptive image feature extraction model according to an embodiment of the present invention; Figure 3 It is a flowchart of neighboring image retrieval according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] The present invention will be further clarified below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent forms of modification of the present invention by those skilled in the art all fall within the scope defined by the appended claims of this application.
[0022] The following embodiments are described by taking the application of using images and text descriptions to complete target image retrieval in the industrial quality inspection scenario as a specific example. The image-text pairs processed by the method have typical industrial feature dimensions. Taking semiconductor wafer defect detection as an example, each high-resolution microscopic image corresponds to structured inspection report text, which includes defect types (such as scratches, particle contamination, lithography anomalies, etc.), defect locations (accurately marked in the wafer coordinate system), size parameters (micrometer-level measurement data), and severity levels (critical defects, major defects, minor defects). By constructing a multi-modal retrieval model based on deep learning, the method can achieve intelligent matching between defect images and the historical case library, provide reference for the disposal solutions of similar defects for quality inspection engineers, and significantly improve the efficiency and accuracy of defect analysis. This application scenario puts forward higher requirements for the fine-grained alignment of image and text features, and requires the model to accurately understand the complex mapping relationship between the microscopic features in the microscopic image and professional inspection terms.
[0023] As Figure 1 To complete the data collection process, first clarify that the input data includes the query image and the concerned conditional text. To ensure the quality consistency of multi-modal data, the method implements a strict three-level data filtering mechanism: First, eliminate obviously unqualified samples such as low-resolution and incomplete text through an automated cleaning pipeline; then use a pre-trained multi-modal model for semantic matching verification, and only retain high-quality data pairs with a text-image similarity exceeding 0.85; finally, conduct sampling audits by domain experts to ensure the accuracy of the annotation of key attributes. This quality control process enables the final constructed training set to maintain a high degree of semantic consistency in different fields, laying a reliable data foundation for subsequent graph neural network modeling. The images obtained after this filtering step usually reach the tens of millions level, and the covered field categories and feature distributions are extremely wide. To effectively process this large-scale and high-dimensional data, a hierarchical clustering algorithm is used for processing. First, use the PCA dimensionality reduction technique to compress the image features from the original dimension to 256 dimensions, which greatly improves the calculation efficiency while ensuring more than 95% of the feature information. Subsequently, an improved K-means++ algorithm is applied for clustering analysis, and at the same time, different clusters are sampled, that is, the intra-class diversity index, to ensure that the features of the sampled data distribution can naturally belong to different clustering clusters, and complete the multi-grained semantic clustering of the images. During the clustering process, a balance constraint condition is specially designed, that is, by setting a threshold to ensure that the sampling is as evenly distributed as possible in as many clusters as possible, that is, the inter-class discrimination index. Prevent the situation where some clusters have too many samples and other clusters have too few samples. After clustering, the system randomly selects representative samples from each cluster in proportion, and finally obtains a training sample set with a moderate quantity but a balanced distribution.
[0024] After the dataset collection is completed, the construction of a text-adaptive image feature extraction model can autonomously adjust the focus of analyzing images according to different detection requirements. Its core technical architecture is as Figure 2 shown. This solution first extracts multi-scale features of microscopic images based on large-scale or industrial-level pre-trained models: the high-resolution detection images are divided into several detection unit regions, and each region is encoded into a feature vector with clear physical meaning through a deep network, that is, the image patch features are extracted; at the same time, text information such as standard process documents and defect description reports is converted into structured semantic features through a domain-specific language model. Based on this, the system constructs a heterogeneous feature map for industrial quality inspection, in which the patch features of the images in the test samples maintain the actual physical space arrangement relationship, and the process text features are introduced as quality determination nodes to establish an interpretable connection relationship with each detection unit, establishing a multi-modal heterogeneous graph. On this heterogeneous graph, the system uses a multi-layer graph attention neural network module for feature extraction and learning. Through the adaptive calculation of the attention mechanism, it can dynamically adjust the attention weights to different image regions according to the semantic content of the conditional text. For example, when processing the detection task of "locating all particle contaminants larger than 5 microns", the system will automatically enhance the attention weight to the pollution-sensitive regions; while when executing the instruction of "identifying the uneven photoresist coating defect", it will focus on analyzing the texture features of the coating region. This intelligent feature focusing mechanism enables the system to accurately capture the key features most relevant to the quality inspection requirements, achieving a dual optimization effect of improving the defect detection rate and reducing the false positive rate in actual production line applications.
[0025] Finally, after the image feature extraction enhanced by text conditions in the previous step, the system further adopts a refined retrieval strategy based on the graph structure to improve the result quality and complete the test. As Figure 3The retrieval of neighboring images and the sample testing process for the target image are completed as shown. First, based on the constructed image semantic relationship graph, the system uses an optimized approximate nearest neighbor search algorithm, such as measuring distances according to Euclidean distance or cosine similarity, to quickly locate the k candidate image nodes that are most similar to the query features. This process not only considers the global feature similarity but also particularly focuses on the matching degree between the local regions of the image and the text conditions. Subsequently, the system innovatively designs a multi-channel feature aggregation mechanism. On the one hand, it retains the most relevant semantic information through attention-weighted feature fusion. On the other hand, it uses multi-scale pooling operations to capture important features at different granularities. At the same time, according to the characteristics of the deep neural network training process, by adding the input of the previous layer network to the next layer network, that is, introducing residual connections to ensure the integrity of the original features. During the actual retrieval process, the system will use a gating mechanism to dynamically adjust the aggregation weights. The test examples input by the user obtain the final enhanced features through the above process, and the target image is selected from the candidate images using the features. This compound feature enhancement method can effectively integrate visual similarity and semantic relevance, and finally generate a more discriminative target image representation to complete the retrieval task. It shows significant advantages when dealing with complex compound queries and provides a reliable solution for scenarios such as e-commerce platforms and medical image retrieval that require high-precision cross-modal search.
[0026] Obviously, those skilled in the art should understand that each step of the above-described method of the graphic retrieval graph network based on model reuse in the embodiments of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to be implemented. In this way, the embodiments of the present invention are not limited to any specific combination of hardware and software.
Claims
1. A graphic retrieval graph network method based on model reuse, characterized in that It includes three steps: image text data acquisition, construction of a text-adaptive image feature extraction model, and neighboring image retrieval. First, collect image data with different distributions and categories from the image text dataset in the public data source. Each image is accompanied by semantically related captions, tags, or descriptive texts. Build a multi-level data preprocessing pipeline to denoise and filter the data in the original image text dataset. Then, construct a text-adaptive image feature extraction model. This step aims to selectively extract relevant features of the image based on the input text conditions. The specific process of constructing the text-adaptive image feature extraction model is to first use a pre-trained model to extract the meta-features of the image and text, construct the initial node representation, design the topological structure of a single image, where the segmented regions of the image form homogeneous subgraphs, and the cross-modal edges are dynamically generated through the graph attention mechanism; generate condition text-adaptive image features. Finally, the specific process of retrieving neighboring images is to use the image nodes and text nodes to form homogeneous subgraphs respectively, and based on the graph sampling and clustering algorithm, obtain similar images of the query image to supplement the context information and generate the target image features to complete the retrieval.
2. The graph network method for image-text retrieval based on model reuse according to claim 1, wherein The specific steps of the image text dataset acquisition are as follows: Step 100, construct an image text data acquisition system, obtain the original image information from the public data source, and record the meta-information of the image at the same time; Step 101, obtain the labeled text of the image at the same network source address as the original image information, and synchronously obtain multi-dimensional labeled texts related to the image from the same network source address based on the semantic association rule; Step 102, adopt a three-level data cleaning and filtering mechanism, namely automatic data extraction, filtering of semantic relations by the pre-trained model, and expert review; the specific mechanism of filtering semantic relations by the pre-trained model is to deploy a pre-trained multi-modal model based on the Transformer architecture, realize automatic semantic matching by calculating the cross-modal similarity score of the image-text pair, and set a threshold to complete the preliminary filtering of low-quality samples; Step 103, adopt a hierarchical clustering algorithm to perform multi-granularity semantic clustering on the image dataset, and optimize the sampling based on the intra-class diversity and inter-class discrimination metrics to ensure the balance and representativeness of the final dataset in both the visual semantics and text description dimensions; After the above steps 100 - 103, an effective training image-text dataset is obtained.
3. The method for retrieving graphic network based on model reuse according to claim 1, wherein The specific steps of the process of constructing the text-adaptive image feature extraction model are as follows: Step 200, use an industrially pre-trained model to extract the block features and text features of the images in the image-text dataset. The block features are non-overlapping image blocks segmented into several parts, providing more fine-grained image meta-features for the adaptive image feature extraction model; Step 201, construct an image block topology graph based on the spatial relationship, where each image block serves as a graph node, establish an edge relationship based on the 4-neighborhood connection strategy to form a regular grid graph structure; Step 202, introduce the text features as global context nodes into the graph structure, establish cross-modal associations with all image block nodes through a fully connected manner, and construct a heterogeneous base graph; Step 203: Adopt a multi-head graph attention mechanism, set the dual-constraint contrastive learning loss function of unimodal and bimodal as the training objective, and complete the training with the intra-modal contrast loss of the image and the image-text bimodal contrast loss as the objectives; based on the heterogeneous base graph generated in Step 202, complete the parameter update of the multi-head graph attention module to generate an image feature representation adapted to the text condition.
4. The graphic retrieval graph network method based on model reuse according to claim 1, characterized in that, The steps for retrieving adjacent images based on the image features are specifically as follows: Step 300: Construct an image semantic relationship graph based on feature similarity, and use cosine similarity to measure the semantic association strength of the fused features; Step 301: Adopt an importance-based random walk sampling algorithm to obtain the k-nearest neighbors of the query image starting from the query node; Step 302: Design a multi-channel feature aggregation module, namely, a neighboring feature attention weighted aggregation and a maximum pooling to capture significant feature aggregation module, and finally dynamically fuse the features extracted by weighted aggregation and the features captured by maximum pooling through a gating mechanism; Step 303: Construct a dataset required for sample evaluation, including: 1) a candidate image library; 2) a reference image set; 3) multi-modal query conditions; Step 304: Use an approximate nearest neighbor search algorithm to quickly match similar images in the feature library.
5. A computer device, characterized in that: The computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the above computer program, it implements the steps of the graph network method for text and image retrieval based on model reuse according to any one of claims 1-4.
6. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program for executing the graph network method for text and image retrieval based on model reuse according to any one of claims 1-4.
Citation Information
Patent Citations
Unsupervised cross-modal hash retrieval method based on noisy label learning
CN112836068A
Combined commodity retrieval method and system based on multi-modal pre-training model
CN114840705A
Cross-modal image-text retrieval method and system based on image-text semantic similarity optimization
CN118484545A
Commodity information processing and querying method and system
CN119377433A
Image-text retrieval method based on comparative learning and modal fusion
CN119441512A
Cited By
Photoetching process window detection method based on contrast learning
CN116482943A
A lithography process window detection method based on contrast learning
CN116482943B