A fabric cross-modal image and text retrieval method based on multi-level representation

Through the cross-modal graphic and text search method of fabric products based on multi-level characterization, the problem of mutual graphic and text search of fabric products is solved, and the cross-modal mutual graphic and text data is realized, which improves the digitalization and intelligence level of manufacturing industry.

CN115168634BActive Publication Date: 2025-06-06JIANGNAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210922659.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-02
Publication Date
2025-06-06
Estimated Expiration
2042-08-02

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively solve the problem of mutual inspection of fabric products with pictures and texts, especially when fabric products are difficult to subdivided, manual labeling is time-consuming and labor-intensive, and keyword subjective.

Method used

A cross-modal graphic and text search method based on multi-level representation is proposed. By constructing a multi-level image representation model and a multi-level text representation model, combined with a bidirectional masking repair model, the hierarchical matching and similarity measurement of graphic and text features are achieved.

Benefits of technology

It realizes cross-modal mutual inspection of fabric images and text data, meets the flexible search needs of different users, improves the design, production and operation efficiency in flexible manufacturing, and promotes the digital and intelligent transformation of the manufacturing industry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115168634B_ABST
    Figure CN115168634B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of fabric retrieval methods, and relates to a fabric cross-modal image-text retrieval method based on multi-level representation. The method steps are as follows: establish a product library containing image and text data; construct an image multi-level representation model to process images; construct a text multi-level representation model to process text, obtain a multi-level feature description of text data in the product library, and form a corresponding relationship with the multi-level feature description of image data; construct a graphic-text hierarchical feature matching model, process the obtained graphic-text multi-level feature description, and perform hierarchical matching of graphic-text features; formulate a retrieval strategy, measure the similarity of graphic-text features, and display the corresponding text or image in order according to the size of the similarity; call out the fabric process sheet corresponding to the image or the image corresponding to the text in the retrieval result to guide production. The present invention has high retrieval accuracy and flexibility, and has great potential in the industrial application field of cross-modal retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of fabric retrieval methods and relates to a fabric cross-modal image-text retrieval method based on multi-level representation. Background Art

[0002] The increasing consumption level has led to the rapid changes in fabric styles and styles. In order to adapt to the changes in the fabric market, fabric manufacturers have gradually turned to a small-batch, multi-variety production model. The rapid replacement of fabric products under this model makes it difficult for companies to query existing product information and it is difficult to give full play to the advantages of historical production experience. Content-based image retrieval has solved the problem of fabric query difficulties to a certain extent, but it is difficult to cope with the two major needs of text query intent image and image query text process sheet. Text-based image retrieval can solve the former need, but fabric products are usually difficult to segment, manual labeling is time-consuming and laborious, and keywords are highly subjective. With the development of multi-source heterogeneous data, the mutual query between fabric images and text has become a problem that textile companies need to solve urgently. Cross-modal image and text retrieval technology can quickly obtain the corresponding text description or intent image by establishing a matching relationship between image and text features, which has important research value for solving the problem of mutual query between fabric product images and text.

[0003] At present, there are no reports on cross-modal retrieval of fabrics. The existing general cross-modal image-text retrieval does not take into account the characteristics of fabric products. Its representation method is difficult to fully represent the hierarchical information of fabric images with strong heterogeneity, and it can deal with the situation where some information of fabric image-text modalities is missing. By establishing a cross-modal image-text retrieval method for fabrics based on multi-level representation, it can meet the retrieval needs of fabric images or texts as query conditions, improve the flexibility of fabric retrieval, and quickly obtain the required text process sheets or intended images. Summary of the invention

[0004] The purpose of the present invention is to propose an efficient, accurate and robust cross-modal fabric image and text retrieval method based on multi-level representation, which can flexibly retrieve intended images or product process sheets for guiding production.

[0005] Based on the above purpose, the present invention provides a fabric cross-modal image and text retrieval method based on multi-level representation, comprising the following steps:

[0006] S1: Build a product library containing image and text data;

[0007] Paired image and text data are selected from the product library to construct a cross-modal image-text retrieval dataset for model training and verification, which mainly includes a training set, a verification set, and a test set.

[0008] S2: Construct a multi-level image representation model to process the image and obtain a multi-level feature description of the image data in the product library;

[0009] The multi-level image representation model uses a convolutional neural network as the underlying framework, builds a multi-task image classification model from multiple perspectives, and mines features at different levels of the image.

[0010] S3: Construct a multi-level text representation model to process text, obtain a multi-level feature description of the text data in the product library, and form a corresponding relationship with the multi-level feature description of the image data;

[0011] The multi-level text representation model uses a bidirectional recurrent neural network as the underlying framework, combines the attention mechanism to extract text keywords to simplify complex semantic dependency information, and adds global constraints for hierarchical representation.

[0012] S4: Build a hierarchical feature matching model for images and texts, process the multi-level feature descriptions of images and texts obtained in S2 and S3, and perform hierarchical matching of image and text features;

[0013] The image-text level feature matching model matches image-text features at different levels by designing a bidirectional masking repair model, and constrains global similarity in a joint embedding space, thereby reducing the granularity of image-text matching and further bridging the heterogeneous differences between images and texts.

[0014] S5: Formulate a retrieval strategy, measure the similarity of image and text features, and display the corresponding text or image in order according to the similarity;

[0015] The retrieval strategy divides the data in the product library into retrieval pools according to the hierarchical category predictions of the image and text multi-level representation model constructed by S2 and S3, refines the search space step by step, determines the retrieval scenario according to the category distribution probability, and judges whether to search across pools and the number of cross-pools.

[0016] S6: The product process sheet or the image corresponding to the text in the search result is retrieved to guide production.

[0017] The product process sheet includes product title, description and attribute information.

[0018] Beneficial effects of the present invention:

[0019] Starting from the retrieval needs of fabric production enterprises, the present invention proposes a fabric cross-modal image and text retrieval method based on multi-level representation. Based on the hierarchical characteristics within the fabric image and text information modality and the strong heterogeneity between modalities, a fabric image and text representation model corresponding to the hierarchical features is constructed to fully express the hierarchical information of the image and text data. By constructing a hierarchical feature matching model for images and texts, the hierarchical matching of image and text features is realized using the idea of ​​bidirectional masking and repair, so as to facilitate the subsequent measurement of image and text feature similarity. A cross-modal image and text retrieval strategy is formulated, a retrieval pool is constructed and it is determined whether to search across the pool, and the similarity of image and text features is measured to solve the problem of missing partial modal information of fabrics. Cross-modal mutual query of fabric images and text data can meet the flexible retrieval needs of different users, improve the design, production and operation efficiency in flexible manufacturing, and thus promote the digital and intelligent transformation of the manufacturing industry. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 This is a flow chart of a method for cross-modal image and text retrieval of fabrics based on multi-level representation according to a preferred embodiment of the present invention.

[0021] Figure 2 are paired image and text data.

[0022] Figure 3 A multi-level representation model for images.

[0023] Figure 4 It is a picture-text level feature matching model.

[0024] Figure 5 shows an example of cross-modal image-text retrieval. (a) is a text query image, and (b) is an image query text. DETAILED DESCRIPTION

[0025] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings.

[0026] The embodiment of the present invention provides a fabric cross-modal image-text retrieval method based on multi-level representation, comprising the following steps:

[0027] S1: Build a product library containing image and text data;

[0028] S2: Construct a multi-level image representation model to process the image and obtain a multi-level feature description of the image data in the product library;

[0029] S3: Construct a multi-level text representation model to process text, obtain a multi-level feature description of the text data in the product library, and form a corresponding relationship with the multi-level feature description of the image data;

[0030] S4: Build a hierarchical feature matching model for images and texts, process the multi-level feature descriptions of images and texts obtained in S2 and S3, and perform hierarchical matching of image and text features;

[0031] S5: Formulate a retrieval strategy, measure the similarity of image and text features, and display the corresponding text or image in order according to the similarity;

[0032] S6: The fabric process sheet or the image corresponding to the text in the search result is retrieved to guide production.

[0033] In order to explain the specific implementation of the present invention in detail, the present invention collects more than 80,000 fabric images and text data from fabric production enterprises as a product database, and selects corresponding image pairs to construct a cross-modal image-text retrieval dataset, and the retrieval performance is better than the existing cross-modal image-text retrieval method. As a preferred embodiment, refer to Figure 1 , which is a flow chart of a fabric cross-modal graphic and text retrieval method based on multi-level representation according to a preferred embodiment of the present invention.

[0034] The method of this embodiment includes the following steps:

[0035] Step S1: Create a product library containing image and text data.

[0036] In this step, paired image and text data are selected from the product library to construct a cross-modal image-text retrieval dataset for model training and verification, which mainly includes a training set, a verification set, and a test set. Figure 2 .

[0037] Step S2: construct a multi-level image representation model to process the image and obtain a multi-level feature description of the image data in the product library.

[0038] In this step, the constructed multi-level image representation model uses a convolutional neural network as the underlying structure, and constructs a multi-task classification model from multiple perspectives to guide the learning of multi-level feature descriptions of images.

[0039] Furthermore, this embodiment uses the VGG-16 network as the underlying structure and constructs a model from five perspectives: fabric pattern, organization, style, color, and category. Figure 3 Taking the fabric characterization model of two tasks as an example, the loss function designed by the present invention is defined as follows:

[0040]

[0041] in, and represents the cross entropy loss function, {W,s 1 ,s 2} are network learning parameters.

[0042] Step S3: construct a text multi-level representation model to process the text, obtain a multi-level feature description of the text data in the product library, and form a corresponding relationship with the multi-level feature description of the image data;

[0043] In this step, the constructed multi-level text representation model uses a bidirectional convolutional neural network as the underlying structure, combines the attention mechanism to extract text keywords to simplify complex semantic dependency information, and adds global constraints for hierarchical representation.

[0044] Furthermore, this embodiment uses a bidirectional long short-term memory network (bi-LSTM) as the underlying structure, whose hidden layer output at the nth word is V, and the word vector is obtained by word-level pooling operation: In the text category attention module, the Hadamard product is used to obtain the previous level information ω h-1 Features Assumptions Represents the weight matrix, using the feature representation S of the category layer h h Execute different categories | C h |Attention, get the text category attention matrix Get the feature representation of the associated text category Assumptions and are the weight matrix and bias respectively, represents a nonlinear activation function, then the feature representation of layer h is A h As shown in the following formula.

[0045]

[0046] For global features It can be obtained by aggregating the features of all layers through hierarchical pooling operations.

[0047] Step S4: constructing a hierarchical feature matching model for images and texts, processing the multi-level feature descriptions of images and texts obtained in S2 and S3, and performing hierarchical matching of image and text features;

[0048] In this step, the constructed image-text level feature matching model refers to Figure 4 A bidirectional masking and inpainting model is designed to match image and text features at different levels, and global similarity is constrained in the joint embedding space. Each time, the features of a certain level of image or text features are masked and inpainted using the corresponding text or image features, thus achieving image and text level feature matching.

[0049] Furthermore, the global constraint maps the image-text features I and T to the joint embedding space so that the difference between the similarity of matching image-text pairs and the similarity of non-matching image-text pairs is as large as possible. as the global optimization goal.

[0050]

[0051] Among them, d(.) represents the similarity measurement function, α represents the margin parameter, [x] + =max(x,0). (I,T) represents a matching image-text pair, and (I′,T) and (I,T′) represent non-matching image-text pairs.

[0052] For the bidirectional masking inpainting model, it is assumed that the inpainted image and text feature vectors are and The feature dimension is D, then the loss function of image and text masking restoration is and The design is as follows:

[0053]

[0054] Among them, λ is a hyperparameter, M is a binary mask, 0 represents the masked part and 1 represents the original part.

[0055] The model is trained by combining the loss functions of global matching and hierarchical matching, and the corresponding weights β are set 1 , β 2 and β 3 , and obtain the final objective function

[0056]

[0057] Step S5: formulate a retrieval strategy, measure the similarity of image and text features, and display the corresponding text or image in order according to the similarity;

[0058] In this step, the retrieval strategy divides the data in the product library into retrieval pools according to the hierarchical category predictions of the image and text multi-level representation model constructed by S2 and S3, refines the search space step by step, determines the retrieval scenario according to the category distribution probability, and judges whether to search across pools and the number of cross-pools.

[0059] Assume that the top three category distribution probabilities of the model output are P 1 , P 2 and P 3 , set P 2 / P 1 and P 3 / P 1Indicates the difference between the query image or text and other categories of images or text, which is used to determine whether to cross-pool retrieval and the number of cross-pool retrieval scenarios. Given different retrieval scenarios R s The threshold λ 1 and λ 2 , R s is defined as follows:

[0060]

[0061] The fabric cross-modal image-text retrieval example of this embodiment is shown in FIG5 . For text retrieval images, given the fabric text to be queried, hierarchical concept phrases W are extracted according to the multi-level representation model of the fabric text. n , and obtain dependency information from the semantic dependency information library to extract text features T n , obtain the fragment feature I of the corresponding category of the image in the retrieval pool n , measures the similarity S between the text feature and all image features in the pool g =d(T g ,I g ). Set the weight α 1 , α 2 and α n The weight of the expression level features is combined to form the final similarity S by combining the similarities of each segment. ti =α 1 S 1 +α 2 S 2 +...+α n S n For image retrieval text, multi-classification is performed based on the constructed multi-level representation model of fabric images, and image features are measured in the retrieval pool. With text features The hierarchical similarity and the global similarity S G =d(I Q ,T P ), where h represents the number of levels, and is expressed by the weight γ h and γ to form the final similarity S it =γ h S h +γS G .

[0062] S6: The fabric process sheet or the image corresponding to the text in the search result is retrieved to guide production.

[0063] In this step, the product process sheet includes product title, description and attribute information.

[0064] Those skilled in the art should understand that the above description is only a specific embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A fabric cross-modal image and text retrieval method based on multi-level representation, It is characterized in that The following steps are involved: S1: Build a product library containing image and text data; Selecting paired image and text data from the product library to construct a cross-modal image-text retrieval dataset for model training and verification, including a training set, a verification set, and a test set; S2: Construct a multi-level image representation model to process the image and obtain a multi-level feature description of the image data in the product library; The multi-level image representation model uses a convolutional neural network as the underlying framework, builds a multi-task image classification model from multiple perspectives, and mines features at different levels of the image; S3: Construct a multi-level text representation model to process text, obtain a multi-level feature description of the text data in the product library, and form a corresponding relationship with the multi-level feature description of the image data; The multi-level text representation model uses a bidirectional recurrent neural network as the underlying framework, extracts text keywords in combination with an attention mechanism to simplify complex semantic dependency information, and adds global constraints for hierarchical representation; S4: Build a hierarchical feature matching model for images and texts, process the multi-level feature descriptions of images and texts obtained in S2 and S3, and perform hierarchical matching of image and text features; In this step, the constructed image-text level feature matching model matches image-text features at different levels by designing a bidirectional masking and repairing model, and constrains global similarity in the joint embedding space. Each time, the features at a certain level of the image or text features are masked and repaired using the corresponding text or image features to achieve image-text level feature matching; The global constraint maps the image-text features I and T to the joint embedding space, so that the difference between the similarity of matching image-text pairs and the similarity of non-matching image-text pairs is as large as possible; the triple loss function is used As a global optimization goal; Where d(.) represents the similarity measure function, α represents the maxgin parameter, [x] + =max(x,0); (I,T) represents a matching image-text pair, (I′,T) and (I,T′) represent a non-matching image-text pair; For the bidirectional masking inpainting model, it is assumed that the inpainted image and text feature vectors are and The feature dimension is D, then the loss function of image and text masking restoration is and The design is as follows: Among them, λ is a hyperparameter, M is a binary mask, 0 represents the masked part, and 1 represents the original part; The model is trained by combining the loss functions of global matching and hierarchical matching, and the corresponding weights β are set 1 , β 2 and β 3 , and obtain the final objective function S5: Formulate a retrieval strategy, measure the similarity of image and text features, and display the corresponding text or image in order according to the similarity; The retrieval strategy divides the data in the product library into retrieval pools according to the hierarchical category prediction of the image and text multi-level representation model constructed by S2 and S3, refines the search space step by step, and determines the retrieval scenario according to the category distribution probability, and determines whether to search across pools and the number of cross-pools; S6: calling out the product process sheet or the image corresponding to the text in the search result to guide production; In the step S2, the constructed multi-level image representation model uses a convolutional neural network as the underlying structure, and constructs a multi-task classification model from multiple perspectives to guide the learning of multi-level feature descriptions of images; The VGG-16 network is selected as the underlying structure, and the model is constructed from five perspectives: fabric pattern, structure, style, color, and category. The loss function designed for the fabric characterization model of the two tasks is defined as follows: in, and represents the cross entropy loss function, {w,s 1 ,s 2 } are network learning parameters.

2. The fabric cross-modal image-text retrieval method based on multi-level representation as claimed in claim 1, It is characterized in that In the step S3, the constructed text multi-level representation model uses a bidirectional convolutional neural network as the underlying structure, extracts text keywords in combination with an attention mechanism to simplify complex semantic dependency information, and adds global constraints for hierarchical representation; The bidirectional long short-term memory network is selected as the underlying structure. Its hidden layer output of the nth word is V. The word vector is obtained through word-level pooling operation: In the text category attention module, the Hadamard product is used to obtain the previous level information ω h-1 Features Assume W s h Represents the weight matrix, using the feature representation S of the category layer h h Execute different categories | C h |Attention, get the text category attention matrix Get the feature representation of the associated text category Assumptions and are the weight matrix and bias respectively, represents a nonlinear activation function, then the feature representation of layer h is A h As shown in the following formula; For global features It can be obtained by aggregating the features of all layers through hierarchical pooling operations.

3. The fabric cross-modal image-text retrieval method based on multi-level representation as claimed in claim 1 or 2, It is characterized in that In the step S5, the retrieval strategy divides the data in the product library into retrieval pools according to the hierarchical category prediction of the image and text multi-level representation model constructed in S2 and S3, refines the search space step by step, and determines the retrieval scenario according to the category distribution probability, and determines whether to search across pools and the number of cross-pools; Assume that the distribution probabilities of the top three categories output by the model are P 1 , P 2 and P 3 , set P 2 / P 1 and P 3 / P 1 Indicates the difference between the query image or text and other categories of images or text, which is used to determine whether to cross-pool retrieval and the number of cross-pool retrieval scenarios; given different retrieval scenarios R s The threshold λ 1 and λ 2 , R s is defined as follows:

4. The fabric cross-modal image-text retrieval method based on multi-level representation as claimed in claim 1 or 2, It is characterized in that In the step S5, for the text retrieval image, given the fabric text to be queried, hierarchical concept phrases W are extracted according to the multi-level representation model of the fabric text. n , and obtain dependency information from the semantic dependency information library to extract text features T n , obtain the fragment feature I of the corresponding category of the image in the retrieval pool n , measures the similarity S between the text feature and all image features in the pool g =d(T g ,I g ); Set weight α 1 , α 2 and α n The weight of the expression level features is combined to form the final similarity S by combining the similarities of each segment. ti =α 1 S 1 + α 2 S 2 +...+α n S n ; For image retrieval text, multi-classification is performed based on the constructed multi-level representation model of fabric images, and image features are measured in the retrieval pool With text features The hierarchical similarity and the global similarity S G =d(I Q ,T P ), where h represents the number of levels, and is expressed by the weight γ h and γ to form the final similarity S it =γ h S h +γS G .

5. The fabric cross-modal image-text retrieval method based on multi-level representation as claimed in claim 3, It is characterized in that In the step S5, for the text retrieval image, given the fabric text to be queried, hierarchical concept phrases W are extracted according to the multi-level representation model of the fabric text. n , and obtain dependency information from the semantic dependency information library to extract text features T n , obtain the fragment feature I of the corresponding category of the image in the retrieval pool n , measures the similarity S between the text feature and all image features in the pool g =d(T g ,I g ); Set weight α 1 , α 2 and α n The weight of the expression level features is combined to form the final similarity S by combining the similarities of each segment. ti =α 1 S 1 + α 2 S 2 +...+α n S n ; For image retrieval text, multi-classification is performed based on the constructed multi-level representation model of fabric images, and image features are measured in the retrieval pool With text features The hierarchical similarity and the global similarity S G =d(I Q ,T P ), where h represents the number of levels, and is expressed by the weight γ h and γ to form the final similarity S it =γ h S h +γS G .

Citation Information

Patent Citations

  • Cross-modal image-text retrieval method

    CN110457516A

  • Cross-modal retrieval method for image-text

    CN114461836A