Cross-modal patent image retrieval method, apparatus, device, and storage medium

By preprocessing and feature extraction of images and text in patent data, and combining feature fusion with GCN network, the problem of low retrieval accuracy caused by ignoring text information in the prior art is solved, and more efficient cross-modal patent image retrieval is achieved.

WO2025140145A1PCT designated stage expired Publication Date: 2025-07-03BEIJING AUGUST MELON TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/141627
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-29
Filing Date
2024-12-23
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

The prior art ignores the text information contained in the image in cross-modal patent image retrieval, resulting in low retrieval accuracy.

Method used

By preprocessing the image and text data in the patent data, the text information in the image is extracted, and feature extraction is combined with Mask R-CNN and BERT embedding, a symmetric adjacency matrix is ​​constructed, and a GCN network is used for feature fusion, a cross-modal hash retrieval method is constructed, and the Hamming distance is calculated to obtain the query result.

Benefits of technology

It improves the accuracy of image retrieval, makes full use of the feature information of images and text, and improves the search accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024141627_03072025_PF_FP_ABST
    Figure CN2024141627_03072025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to a cross-modal patent image retrieval method, an apparatus, a device, and a storage medium. The method comprises: acquiring a query instruction and updated patent data; obtaining patent fusion feature data on the basis of the updated patent data; on the basis of the patent fusion feature data and fusion feature data pre-stored in a preset database, obtaining integrated database fusion feature data, wherein the preset database fusion feature data is obtained by fusing image features and text features in the patent data; determining the type of the query instruction on the basis of the query instruction, and performing calculation to obtain query feature data; and on the basis of the query feature data and the integrated database fusion feature data, performing Hamming calculation to obtain a query result.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-modal patent image retrieval method, device, equipment and storage medium

[0001] This application is filed with the Patent Office of China on December 29, 2023, with application number 2023118734278 and the name of the invention being “Cross-modal patent image retrieval method, apparatus, device and storage medium”, all contents of which are incorporated by reference in this application. Technical Field

[0002] The present application relates to the field of image retrieval, and specifically to a cross-modal patent image retrieval method, apparatus, device and storage medium. Background Art

[0003] Due to the rapid development of deep learning and the high efficiency and low cost of hashing algorithms, deep cross-modal hashing retrieval has made great progress in recent years.

[0004] Cross-modal hashing methods combine text and content for integrated research. Leveraging their respective strengths, they promote efficient and simple image retrieval. Especially in online environments, they combine feature analysis of the web document containing the image to infer its characteristics. Combined with content analysis, they achieve joint indexing for image analysis and retrieval.

[0005] However, most existing methods only focus on image features and ignore the text information contained in the image, resulting in low retrieval accuracy for images containing relatively important text information. Summary of the Invention

[0006] In response to the above problems, the purpose of this application is to provide a cross-modal patent image retrieval method, device, equipment and storage medium.

[0007] According to a first aspect of an embodiment of the present application, a cross-modal patent image retrieval method is provided, characterized in that it includes: obtaining a query instruction and updated patent data; obtaining patent fusion feature data based on the updated patent data; obtaining comprehensive database fusion feature data based on the patent fusion feature data and fusion feature data pre-stored in a preset database; wherein the preset database fusion feature data is obtained by fusing image features and text features in the patent data; judging the type of query instruction based on the query instruction and calculating the query feature data; performing Hamming calculation based on the query feature data and the comprehensive database fusion feature data to obtain a query result.

[0008] Further, it is characterized in that the patent fusion feature data is obtained based on the updated patent data, including: preprocessing the patent image data through the patent image data and patent text data in the updated patent data to obtain text data in the image; obtaining patent splicing text data based on the patent text data and the text data in the image; performing feature extraction on the patent splicing text data and the patent image data respectively to obtain patent image feature data and patent text feature data; and fusing the patent image feature data and the patent text feature data to obtain the patent fusion feature data.

[0009] Furthermore, it is characterized in that the feature extraction is performed on the patent spliced ​​text data and the patent image data respectively to obtain patent image feature data and patent text feature data, including: based on Mask R-CNN, feature extraction is performed on the patent image data to obtain patent image feature data; based on BERT embedding, feature extraction is performed on the patent spliced ​​text data to obtain patent text feature data.

[0010] Further, it is characterized in that the fusion of the patent image feature data and the patent text feature data to obtain the patent fusion feature data includes: constructing a symmetric adjacency matrix based on the patent image feature data and the patent text feature data; and obtaining the patent fusion feature data based on the GCN network based on the patent image feature data, the patent text feature data and the symmetric adjacency matrix.

[0011] Further, it is characterized in that the query instruction type includes one or more of text, image and text and image; according to the query instruction, the query instruction type is judged, and the query feature data is calculated, including: when the query instruction type is text, based on the LSTM network, the query data text feature data is obtained; when the query instruction type is image, according to the query data, the text feature data and the query image feature data in the query data image are obtained, and the query data image feature data are fused; when the query instruction type is text and image, according to the query data, the query image data and the query text data are obtained, the query image data features and the query text splicing data features are calculated, and the fused query feature data is obtained; the query feature data includes one or more of the query data text feature data, the query data image feature data and the fused query feature data.

[0012] According to a second aspect of an embodiment of the present application, a cross-modal patent image retrieval device is provided, characterized in that it includes: a data acquisition module for acquiring query instructions and updated patent data; a patent fusion feature data calculation module for obtaining patent fusion feature data based on the updated patent data; a comprehensive database fusion feature data calculation module for obtaining comprehensive database fusion feature data based on the patent fusion feature data and fusion feature data pre-stored in a preset database; wherein the preset database fusion feature data is obtained by fusing image features and text features in the patent data; a query feature data calculation module for judging the type of query instruction based on the query instruction and calculating the query feature data; a query module for performing Hamming calculation based on the query feature data and the comprehensive database fusion feature data to obtain a query result.

[0013] Furthermore, it is characterized in that the patent fusion feature data calculation module is specifically used to: pre-process the patent image data through the patent image data and patent text data in the updated patent data to obtain the text data in the image; obtain patent splicing text data based on the patent text data and the text data in the image; perform feature extraction on the patent splicing text data and the patent image data respectively to obtain patent image feature data and patent text feature data; fuse the patent image feature data and the patent text feature data to obtain the patent fusion feature data.

[0014] Further, it is characterized in that the query instruction type includes one or more of text, image and text and image; the query feature data calculation module is specifically used to: when the query instruction type is text, obtain the query data text feature data based on the LSTM network; when the query instruction type is image, obtain the text feature data and query image feature data in the query data image according to the query data, and fuse them to obtain the query data image feature data; when the query instruction type is text and image, obtain the query image data and query text data according to the query data, calculate the query image data features and the query text splicing data features, and obtain the fused query feature data; the query feature data includes one or more of the query data text feature data, the query data image feature data and the fused query feature data.

[0015] According to a third aspect of an embodiment of the present application, an electronic device is provided, comprising: a processor; a memory for storing processor-executable instructions; and a processor for reading executable instructions from the memory and executing the instructions to implement a cross-modal patent image retrieval method provided in the first aspect of the present disclosure.

[0016] According to a fourth aspect of an embodiment of the present application, a computer-readable storage medium is provided, on which computer program instructions are stored. When the program instructions are executed by a processor, the steps of a cross-modal patent image retrieval method provided in the first aspect of the present disclosure are implemented.

[0017] The technical solutions provided by the embodiments of the present application may have the following beneficial effects:

[0018] This application obtains text data in the image by pre-processing the database image data, extracting the textual information therein, and combining the characteristics that patent data images may contain more important textual information, thereby increasing the amount of relevant information extracted and utilized, and obtaining text data in the image. The text data in the image is combined with the database text data to obtain database spliced ​​text data, and all text-type data are integrated to facilitate subsequent data extraction. Feature extraction is performed on the database spliced ​​text data and the database image data respectively, making full use of the feature information of different modal data to improve the retrieval accuracy, and the database image feature data and the database text feature data are fused to obtain the database fused feature data, and the two features are sorted and merged to facilitate subsequent retrieval. The query feature data is calculated, and the query result is calculated with the database fused feature data, making full use of the data information of the query data and the database data to obtain high-quality retrieval results. This application fully utilizes the relevant information and improves the image retrieval accuracy by calculating the features of the query data and the database data respectively. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate one or more embodiments of this specification or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0020] FIG1 is a flowchart illustrating a method for cross-modal patent image retrieval according to an exemplary embodiment;

[0021] FIG2 is a schematic diagram of a network structure of a cross-modal patent image retrieval method according to an exemplary embodiment;

[0022] FIG3 is a device diagram showing a cross-modal patent image retrieval method according to an exemplary embodiment. DETAILED DESCRIPTION

[0023] In order to enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below in conjunction with the drawings in one or more embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this document.

[0024] Exemplary Method 1

[0025] The cross-modal patent image retrieval method of this application inputs a query image or query text, and through feature extraction and hash mapping, finds the target image and text in the cross-modal hash features of the database text and database image.

[0026] In step S101, a query instruction and updated patent data are obtained.

[0027] Obtain query instructions, specifically the query instructions entered by the user.

[0028] Obtain updated patent data. Specifically, this application is a cross-modal patent image retrieval method. Patent data will be continuously updated. Based on this feature, a preset time point is set and newly disclosed patent data is obtained after this time point every day.

[0029] In step S102, patent fusion feature data is obtained based on the updated patent data.

[0030] The patent image data is preprocessed using the patent image data and patent text data in the updated patent data to obtain text data in the image.

[0031] There are a large number of illustration images and corresponding illustration descriptions in patent data, some of which contain relatively important text information. Therefore, it is necessary to extract the text in the images.

[0032] Perform grayscale processing on the image to convert the multi-channel color image into a single-channel grayscale image. In this application, the weighted average method is used to obtain the grayscale image. The specific formula is as follows: G(x,y)=0.114B+0.587G+0.299R

[0033] Based on the human eye's observation characteristics, the preset weighting coefficients are 0.114 for blue (B), 0.587 for green (G), and 0.299 for red (R). A Gaussian filter is used to remove noise from the grayscale image to reduce interference. The Sobel operator is used to calculate the horizontal and vertical gradient values ​​of the grayscale image, determining the gradient strength and direction for each pixel. The gradient values ​​of the two adjacent pixels in each pixel's gradient direction are compared, retaining only those pixels with the largest gradient values ​​in that direction. This helps to sharpen image edges. Pixels are classified into three categories: strong edges, weak edges, and non-edges. A higher threshold and a lower threshold are preset. When the gradient value of a pixel exceeds the higher threshold, it is classified as a strong edge. When the gradient value of a pixel is below the lower threshold, it is classified as a non-edge. When the gradient value of a pixel is between the higher and lower thresholds, it is classified as a weak edge. Weak edges are connected to the surrounding strong edges to form a complete edge, ultimately producing the processed grayscale image. The processed grayscale image is binarized, and the vertical edge image is subjected to grayscale cumulative sampling of the pixel values ​​on the x-axis, i.e., aggregation processing. The processing formula is as follows:

[0034] Where E(t,y) is the compression processing function of the edge detection result of the original grayscale image into the x-axis. The parameter in the formula represents the compression factor, y indicates that the processing object is a vertical edge image, the subscript l represents the horizontal compression function of the grayscale image edge detection result in the vertical direction, and α is the step size.

[0035] Through aggregation processing, the space occupied by image data is compressed to increase data stability.

[0036] The corresponding thresholds T1 and T2 are preset to divide the processing results. The formula is as follows:

[0037] Among them, V0 represents the pixel set of the non-text area in the image; H1 represents the pixel set that is currently uncertain, that is, the pixel set that needs to be expanded and extended; H2 can be preliminarily determined to be the pixel set of the text area.

[0038] In this application, the eight-neighborhood method is used for the pixel point set H1, and the preset division threshold is 0.5 for division processing. A 4×4 pixel point set is processed each time, and the pixel point set in H1 is divided into H0 or H2 to complete the image pixel division and obtain the divided pixel point set.

[0039] The pixel set after division is expanded using a 5×5 matrix and the expansion result is calculated using the following formula: i =max(H j ),j∈M

[0040] Among them, j is a pixel point, and M is a set of adjacent pixel point matrices of order 4.

[0041] The expansion result is eroded using a 5×5 matrix and the erosion result is calculated using the following formula: i =min(H j ),j∈M

[0042] Through dilation and erosion processing, noise is eliminated, target features are segmented, and preparation is made for connected area calculation.

[0043] Calculate and analyze the connected areas of corrosion results.

[0044] A connected component is an image region (or blob) consisting of adjacent foreground pixels with the same pixel value. Connected component analysis (Connected Component Labeling) involves identifying and labeling connected regions within an image. Connected component analysis is a common and fundamental method in image analysis and processing. It can be used in any application scenario where foreground objects need to be extracted for subsequent processing. Typically, the object of connected component analysis is a binary image.

[0045] In an embodiment of the present application, based on a two-pass scanning method, a four-neighborhood is used to implement analysis and calculation of connected domains, wherein the two-pass scanning method (Two-Pass) refers to finding and marking all connected domains in an image by scanning the image twice.

[0046] Specifically, the first scan starts from the upper left corner and traverses the pixels, finding the first pixel with a value of 255, label = 1; when the left neighbor and the upper neighbor of the pixel have invalid values, a new label value is set for the pixel, label++, and the record set is recorded; when one of the left neighbor or the upper neighbor of the pixel has a valid value, the label of the valid value pixel is assigned to the label value of the pixel; when the left neighbor and the upper neighbor of the pixel both have valid values, the smaller label value is selected and assigned to the label value of the pixel. The purpose of the second traversal is to merge different digital labels belonging to the same connected domain. Therefore, before the second traversal begins, a union-find (union-find) process is required on the array storing the digital labels. In the union-find structure, the label values ​​belonging to the same connected domain are stored in the same tree structure. The data of the tree structure is stored in a vector or array. The subscript of the vector is the label value, and the stored value of the vector is the label value of the parent node of the label. When the stored value of the vector is 0, it means that the label is the root node. You can use the find algorithm to find the root node of any label. Find the root node of each label. If the root nodes of the two labels are the same, then they belong to the same connected domain. If the root nodes of the two labels are different, then make one root node the parent node of the other, and the two labels belong to the same connected domain.

[0047] In the embodiment of the present application, certain characteristic values ​​of the connected domain are used to filter the connected domain, remove unnecessary connected domains, and determine the scope of the target area.

[0048] Specifically, the features of the connected domain include area, perimeter, etc. This application uses the number of pixels contained in the connected domain as the characteristic value to perform connected domain filtering. The connected domain filtering process first uses the connected domain determination method to divide all target pixels into connected domains, and save each connected domain separately; then set the filtering threshold, compare the size of the connected domain with the threshold, and filter out connected domains that are smaller than (or larger than) the threshold; finally, set all pixels in the filtered connected domain as background pixels, so that the unnecessary connected domains can be deleted and only the required connected domains are retained, that is, the text area in the image is determined. Use a fixed rectangular frame to limit the text area, and use the coordinates of the lower right corner and the upper left corner of the rectangular area to represent it.

[0049] In an embodiment of the present application, after determining the text area in the image, the text area in the image is segmented to obtain text information.

[0050] After positioning, the text area in the image is still a set of pixels as a whole. The set contains not only the stroke information of the text, but also the information of the blank areas between text lines and characters. Therefore, in order to effectively recognize single characters, the characters in the text area need to be segmented. Different text area rows are divided according to the height of the text, and the regional vertical projection calculation is performed on the image text. The character size and the offset in the text line are calculated. The zero-value interval in the result is compared with the preset empirical threshold. The vertical projection calculation result is analyzed to obtain the character size. Text lines with relatively close character heights are grouped into a text line set. Then, each text line set is vertically projected, and the pixel grayscale values ​​of each projection column are accumulated to obtain the segmented data.

[0051] After segmenting the text characters in the text area of ​​the image, normalization is required to process and convert the text of varying sizes within the input image into a uniformly sized character image. OCR technology is then used to perform text recognition on the segmented and normalized data, obtaining the text information within the image. Optical Character Recognition (OCR) is the process of analyzing and identifying image files of text materials to obtain text and layout information. This process essentially recognizes the text within an image and returns it as text.

[0052] Patent spliced ​​text data is obtained based on the patent text data and the text data in the image. The image description text and the text information in the image are combined and stored to obtain a text information set.

[0053] Feature extraction is performed on the patent splicing text data and the patent image data to obtain patent image feature data and patent text feature data.

[0054] As shown in Figure 2, patent image features: input patent image data, use MASK R-CNN as the target detection network, generate n bounding boxes, and the ROI Align layer outputs the feature vector of the bounding box.

[0055] Mask R-CNN consists of two stages. The first stage scans the image and generates proposals (regions that may contain an object). The second stage classifies the proposals and generates bounding boxes and masks. The image is input to Mask R-CNN for data preprocessing. The processed image is then passed to a pretrained neural network to obtain the corresponding feature map. At each convolutional layer, the data is presented in three dimensions, each of which is called a feature map. Regions of interest (ROIs) are defined for each point in the feature map, resulting in multiple ROI candidate boxes.

[0056] ROI (region of interest). In this application's image processing, a region of interest (ROI) is defined as a box around the image being processed. Various operators and functions are used in machine vision software to determine the ROI and perform further image processing.

[0057] These multiple ROI candidate frames are sent to RPN for binary classification and BB regression, some candidate ROIs are filtered out, and the remaining ROIs are subjected to ROI Align operation;

[0058] RPN (Region Proposal Network), the Chinese meaning of Region Proposal is "region selection", which means "extracting candidate boxes". RPN is the network used to extract candidate boxes.

[0059] Patent text features: The text information set obtained after input data preprocessing uses BERT embedding to represent sentence descriptions and obtain a text vector set.

[0060] BERT embedding consists of three parts: token embedding, segment embedding, and position embedding.

[0061] Token Embeddings, vector representations of words; Segment Embeddings, vector representations that help BERT distinguish between the two sentences in a sentence pair. The Segment Embeddings layer has only two vector representations. The first vector assigns 0 to each token in the first sentence, and the second vector assigns 1 to each token in the second sentence. If the input is only a single sentence, then its segment embedding will be all 0s; Position Embeddings, allowing BERT to learn the sequential nature of the input. In the position embedding, the position index of each token in the input sequence. These representations are added element-wise to obtain a synthetic representation of the text vector set.

[0062] Among them, token is the basic unit of model input. In Chinese BERT, a token can be a word or an identifier. In this application, "token" refers to the process of converting each word in the text into a corresponding number or vector representation. Embedding: A dense vector used to represent a token. The token itself is not computable and needs to be mapped to a continuous vector space before subsequent calculations can be performed. The result of this mapping is the embedding corresponding to the token. A fully connected layer is used to convert patent image features and patent text features into vectors of the same dimension d.

[0063] The patent image feature data and the patent text feature data are fused to obtain the patent fusion feature data.

[0064] In the embodiments of the present application, there is a modal gap between image features and text features, which is mainly reflected in the following aspects: Differences in representation: Image features and text features are expressed differently. Image features are usually represented in the form of pixel values, colors, textures, etc., while text features are represented in the organizational structure of words, sentences or paragraphs. Differences in semantic understanding: Image features and text features have different emphases on semantic understanding. Image features mainly focus on visual information, such as objects, scenes, shapes, etc., while text features pay more attention to describing semantic information such as attributes, relationships, emotions, etc. of objects. Differences in data scale: Image data is usually high-dimensional, continuous numerical data, while text data is discrete, symbolic data. Data from different modalities exhibit heterogeneity and variability in their formal structure and distribution. Samples from different modalities are not only difficult to measure directly, but also contain inconsistent semantic information. When mapped to a common semantic space for measurement, this can lead to decreased retrieval performance due to the inclusion of irrelevant information. Therefore, cross-modal hash retrieval, while mapping different modalities to the same common semantic space, must address the semantic gap and establish as much semantic information in the sample as possible that is relevant to the user's search intent. This disparity in data size necessitates the use of different algorithms and techniques during training and processing.

[0065] Therefore, it is necessary to fuse image features and text features so that the similarity between feature vectors is consistent with the correlation between modal data and realize cross-modal tasks.

[0066] Specifically, relevance information in cross-modal retrieval is often unordered. For example, text consists of different word arrangements, while images consist of different pixel arrangements. These differences in underlying representations prevent direct comparison of feature representations across modal data, leading to inability to directly compare and calculate their relevance. Therefore, we propose using a graph structure to capture instance-level associations between candidate sample points in each database. Instance-level associations mean that the more similar two examples are, the more similar their labels are. Each example is considered a point in a graph, and the edges in the graph represent the associations between them. Finally, the representation vector of this graph structure is used as the feature representation of the entire data point. This setup accommodates the inherent unordered nature of associations between examples in cross-modal retrieval scenarios. Specifically, for the i-th data sample point, an undirected graph is constructed, consisting of the points, edges, and corresponding adjacency matrix. To ensure that the similarity between example feature vectors is consistent with the inter-modal data relevance, the database text features and database image features are combined through element-wise multiplication, thereby constructing a symmetric adjacency matrix for the i-th data sample point. This is calculated as follows:

[0067] Among them, G i is the image feature, F i For text features.

[0068] GCN is used to learn the representation vector of the graph structure as the cross-modal expression of each candidate sample point in the database. i As input GCN, GCN contains L layers and uses ReLU as the activation function. The output of each layer can be calculated by layer-by-layer transmission to obtain a cross-modal representation vector, which is passed through a fully connected layer with tanh as the activation function to output a cross-modal hash mapping vector. The cross-modal hash mapping vector is converted into a binary code to facilitate fast image search based on Hamming distance, greatly reducing computational costs and further optimizing search efficiency.

[0069] In step S103, comprehensive database fusion feature data is obtained based on the patent fusion feature data and the fusion feature data pre-stored in the preset database.

[0070] The preset database fusion feature data is obtained by fusing the image features and text features in the patent data using the method in S102. Since the number of patents in the database is large and the features are fixed, the data in the database is pre-featured and fused before being stored in the system.

[0071] In step S104, the query instruction type is determined according to the query instruction, and query feature data is calculated.

[0072] In an embodiment of the present application, a query instruction is used to represent the search information input when searching for patents. Feature extraction is performed on the search instruction to obtain query feature data for subsequent comparison and search with the information in the database. The search instruction is generally a single image or text. When the search information is an image, the text content in the image needs to be considered, but when the search information is text, it is not necessary to extract word vectors and perform feature fusion with the query image. Therefore, the corresponding feature extraction method is slightly different when the search information is an image or text. In an embodiment of the present application, when the query instruction type is text, the query data text feature data is obtained based on the LSTM network; when the query instruction type is an image, the query data image text feature data and the query image feature data are obtained according to the query data, and the query data image feature data are fused to obtain the query data image feature data; when the query instruction type is text and image, the query image data and the query text data are obtained according to the query data, and the query image data features and the query text splicing data features are calculated to obtain the fused query feature data; the query feature data includes one or more of the query data text feature data, the query data image feature data and the fused query feature data. Specifically,

[0073] For the case where the retrieval information is an image: use the method of step S102 to extract text from the query image to obtain query image text information, convert the query text information into text feature data, convert the query image into image feature data, and merge the query text feature data with the query image feature data.

[0074] For the case where the retrieved information is text: LSTM is used to transform the query text, and a fully connected layer is used to map the convolutional features to a hash mapping vector, and then the hash mapping vector is quantized into a binary code through a function.

[0075] The query text is vectorized and input into the first embedding layer, which uses a preset length vector to represent each word. The SpatialDropout1D layer randomly sets the input unit ratio to 0 each time it is updated during training to prevent overfitting. The LSTM layer contains memory units, and the output layer is a fully connected layer with the activation function set to Softmax.

[0076] The Softmax function, or normalized exponential function, is a generalization of the logistic function. It can "compress" an N-dimensional vector z containing any real number into another N-dimensional real vector, so that each element is between (0, 1) and the sum of all elements is 1.

[0077] In the case where the search information is text and images: the method of step S102 is used to perform feature extraction and fusion on the search information to obtain fused query feature data.

[0078] In step S105, a Hamming calculation is performed based on the query feature data and the integrated database fusion feature data to obtain a query result.

[0079] The Hamming distance between the fusion feature data of the comprehensive database and the query feature data is calculated. The smaller the Hamming distance, the higher the similarity. Therefore, patent samples are recommended from small to large according to the Hamming distance value to obtain the search results.

[0080] Exemplary Method 2

[0081] The model is trained. The objective function calculates the error between the feature result and the true label S, and commands the model parameters to be refreshed through the error back propagation algorithm.

[0082] The drawings in the patent application specification are selected. Each drawing has a corresponding caption. The drawings and the caption constitute a sample, from which 10,000 are selected as a training set, 5,000 as a retrieval set, and 5,000 as a test set.

[0083] As shown in Figure 2, the search network of this application consists of two parts: the search network part extracts features from the patent database, and the query network part extracts query features. S102 and S103 are implemented by the search network part, while S104 is implemented by the query network part. The two parts of the search network of this application are trained and optimized alternately.

[0084] Specifically, first, the query hash function in the query network is fixed, the input data is forward-transmitted, the error gradient is back-propagated, the database cross-modal hash code in the retrieval network is optimized, and the parameters in the retrieval network model are updated; then, the database cross-modal hash code in the retrieval network is fixed, the input data is forward-transmitted, the error gradient is back-propagated, the query hash function in the query network is optimized, and the parameters in the query network model are updated. The two steps are optimized alternately, and according to the preset number of iterations, the trained image retrieval network is finally obtained.

[0085] First, the query hash function is fixed, and the training set is input one by one and forward-propagated. The cross-modal hash mapping vector of the patent database, the query text hash mapping vector, and the query image hash mapping vector are calculated. The error gradient is calculated and back-propagated. The parameters of the patent database feature extraction part of the network model are updated accordingly.

[0086] Then fix the cross-modal hash code of the patent database, input and forward conduct the training set one by one, calculate the cross-modal hash mapping vector of the patent database and the query text hash mapping vector, query image hash mapping vector, calculate the error gradient, back-propagate the error gradient, and update the parameters of the query feature extraction part of the network model accordingly.

[0087] Set up an objective function to find the data in the database that is most similar to the input data as the search result.

[0088] The objective function consists of the L2 norm and constraints. The L2 norm measures the difference between the similarity score and the inner product of the hash code. The formula is as follows:

[0089] Among them, S is the cross-modal similarity matrix generated by the supervision information. When the image x i and text y i , (or text y i and image x i ); similarity score s ij =1, otherwise s ij =-1. b i 、b j Binary codes converted from different cross-modal hash map vectors, is the inner product between hash codes.

[0090] in, The Hamming distance is calculated based on the Hamming distance to calculate the similarity between data points. The Hamming distance is calculated by comparing each bit of the vector to see if they are the same. If they are different, the Hamming distance is increased by 1, and the Hamming distance is obtained. The higher the vector similarity, the smaller the corresponding Hamming distance. In this application, the Hamming distance between two hash codes can be expressed as:

[0091] Where K is the hash code length.

[0092] The L2 norm is a function that represents the concept of “distance”. The L2 norm can prevent overfitting and improve the generalization ability of the model.

[0093] Replacing the discrete hash code with a continuous hash mapping vector and introducing a constraint on the modal-dependent hash function, the objective function formula is as follows:

[0094] Among them, u i 、u j are different cross-modal hash mapping vectors, and is the hash map output of the modality-dependent hash function, Hash map output for the query image, Hash map output for query text, u q is the cross-modal hash mapping vector; λ is the weight ratio of the L2 norm and the constraint condition.

[0095] The constraints ensure that the query data points can be projected into the cross-modal hash space and added to the query hash function.

[0096] The L2 norm calculated from the database's cross-modal hash mapping vectors is used as the pairwise similarity loss. The query image and query text hash mapping outputs form the constraints. Together, the pairwise similarity loss and the constraints form the objective function formula.

[0097] During the training process, the objective function, u i 、 and Calculate the gradient and assist in training. The formula is as follows:

[0098] Exemplary devices

[0099] In the embodiment of the present application, as shown in FIG3 , it includes 301 a data acquisition module, 302 a patent fusion feature data calculation module, 303 a comprehensive database fusion feature data calculation module, 304 a query feature data calculation module and 305 a query module.

[0100] The data acquisition module 301 is used to acquire query instructions and updated patent data;

[0101] Patent fusion feature data calculation module 302 obtains patent fusion feature data based on the updated patent data; preprocesses the patent image data using the patent image data and patent text data in the updated patent data to obtain image text data; obtains patent spliced ​​text data based on the patent text data and the image text data; performs feature extraction on the patent spliced ​​text data and the patent image data, respectively, to obtain patent image feature data and patent text feature data; performs feature extraction on the patent image data based on Mask R-CNN to obtain patent image feature data; performs feature extraction on the patent spliced ​​text data based on BERT embedding to obtain patent text feature data. The patent image feature data and the patent text feature data are fused to obtain the patent fusion feature data. A symmetric adjacency matrix is ​​constructed based on the patent image feature data and the patent text feature data; and patent fusion feature data is obtained based on the GCN network based on the patent image feature data, the patent text feature data, and the symmetric adjacency matrix.

[0102] The integrated database fusion feature data calculation module 303 is used to obtain integrated database fusion feature data based on the patent fusion feature data and the fusion feature data pre-stored in the preset database; wherein the preset database fusion feature data is obtained by fusing the image features and text features in the patent data;

[0103] The query feature data calculation module 304 is used to determine the type of query instruction based on the query instruction and calculate the query feature data. The query instruction type includes one or more of text, image, and text and image. When the query instruction type is text, the query data text feature data is obtained based on the LSTM network; when the query instruction type is image, the text feature data in the query data image and the query image feature data are obtained according to the query data, and the query data image feature data are fused together; when the query instruction type is text and image, the query image data and the query text data are obtained according to the query data, the query image data features and the query text splicing data features are calculated, and the fused query feature data is obtained; the query feature data includes one or more of the query data text feature data, the query data image feature data, and the fused query feature data.

[0104] The query module 305 is configured to perform Hamming calculation based on the query feature data and the integrated database fusion feature data to obtain a query result.

[0105] Exemplary electronic devices

[0106] This embodiment proposes an electronic device, comprising: one or more processors, and an internal memory and an external memory, wherein the internal memory stores instructions, and when the instructions are executed by the one or more processors, the one or more processors execute a cross-modal patent image retrieval method described in any of the aforementioned embodiments.

[0107] The processor is configured to execute all or part of the steps of the cross-modal patent image retrieval method described in the embodiment. The memory is configured to store various types of data, such as instructions for any application or method in the electronic device, as well as application-related data.

[0108] The processor can be an application specific integrated circuit (ASIC), a digital signal processor (DSP), a programmable logic device (PLD), a field programmable gate array (FPGA), a controller, a microcontroller, a microprocessor or other electronic components, and is used to execute a cross-modal patent image retrieval method described in the embodiment.

[0109] Computer storage media

[0110] The computer-readable storage medium stores a computer program, and when the computer program is executed by one or more processors, it implements a cross-modal patent image retrieval method as described in any of the aforementioned embodiments.

[0111] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0112] In the 1930s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures like diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements today can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to hire a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly done using "logic compiler" software. This is similar to the software compiler used when developing programs. Before compilation, the original code must also be written in a specific programming language, called a hardware description language (HDL). There is not just one HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that by simply programming the method flow in one of these hardware description languages ​​and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.

[0113] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that in addition to implementing the controller in a purely computer-readable program code format, the controller can be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be considered as structures within the hardware component. Or even, the devices for implementing various functions can be considered as both software modules that implement the method and structures within the hardware component.

[0114] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0115] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing the embodiments of this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0116] Those skilled in the art will appreciate that one or more embodiments of this specification may be provided as a method, system, or computer program product. Thus, one or more embodiments of this specification may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0117] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0118] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0119] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0120] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0121] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0122] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0123] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0124] One or more embodiments of this specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. One or more embodiments of this specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0125] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.

[0126] The foregoing description is merely an example of the present invention and is not intended to limit the present invention. Persons skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be included within the scope of the claims herein.

Claims

1. A cross-modal patent image retrieval method, characterized in that, Including: Obtain a query instruction and updated patent data; Obtain patent fusion feature data according to the updated patent data; Obtain comprehensive database fusion feature data according to the patent fusion feature data and the fusion feature data pre-stored in a preset database; wherein, the preset database fusion feature data is obtained by fusing the image feature and the text feature in the patent data; Judge the type of the query instruction according to the query instruction, and calculate to obtain query feature data; Perform Hamming calculation according to the query feature data and the comprehensive database fusion feature data to obtain a query result.

2. The method according to claim 1, characterized in that The obtaining patent fusion feature data according to the updated patent data includes: Perform preprocessing on the patent image data in the updated patent data through the patent image data and the patent text data in the updated patent data to obtain the text data in the image; Obtain patent spliced text data according to the patent text data and the text data in the image; Perform feature extraction on the patent spliced text data and the patent image data respectively to obtain patent image feature data and patent text feature data; Fuse the patent image feature data and the patent text feature data to obtain the patent fusion feature data.

3. The method according to claim 2, wherein The performing feature extraction on the patent spliced text data and the patent image data respectively to obtain patent image feature data and patent text feature data includes: Perform feature extraction on the patent image data based on Mask R-CNN to obtain patent image feature data; Perform feature extraction on the patent spliced text data based on BERT embedding to obtain patent text feature data.

4. The method according to claim 2, wherein The fusing the patent image feature data and the patent text feature data to obtain the patent fusion feature data includes: Construct a symmetric adjacency matrix according to the patent image feature data and the patent text feature data; Based on the patent image feature data, the patent text feature data and the symmetric adjacency matrix, obtain patent fusion feature data based on a GCN network.

5. The method according to claim 1, wherein The types of the query instruction include one or more of text, image, and text and image; Judging the type of the query instruction according to the query instruction and calculating to obtain query feature data includes: When the type of the query instruction is text, obtain query data text feature data based on an LSTM network; When the type of the query instruction is image, obtain the text feature data in the query data image and the query image feature data according to the query data, and fuse them to obtain query data image feature data; When the type of the query instruction is text and image, obtain query image data and query text data according to the query data, calculate the query image data feature and the query text spliced data feature to obtain fused query feature data; The query feature data includes one or more of query data text feature data, query data image feature data, and fused query feature data.

6. A cross-modal patent image retrieval device, characterized in that, Including: A data acquisition module for acquiring a query instruction and updated patent data; A patent fusion feature data calculation module for obtaining patent fusion feature data according to the updated patent data; The comprehensive database fusion feature data calculation module is used to obtain the comprehensive database fusion feature data according to the patent fusion feature data and the fusion feature data pre-stored in the preset database; wherein, the preset database fusion feature data is obtained by fusing the image features and text features in the patent data; The query feature data calculation module is used to judge the type of the query instruction according to the query instruction and calculate the query feature data; The query module is used to perform Hamming calculation according to the query feature data and the comprehensive database fusion feature data to obtain the query result.

7. The device according to claim 6, characterized in that, The patent fusion feature data calculation module specifically is used for: Preprocessing the patent image data in the updated patent data to obtain the text data in the image; Obtaining the patent spliced text data according to the patent text data and the text data in the image; Performing feature extraction on the patent spliced text data and the patent image data respectively to obtain the patent image feature data and the patent text feature data; Fusing the patent image feature data and the patent text feature data to obtain the patent fusion feature data.

8. The device according to claim 6, characterized in that, The types of the query instructions include one or more of text, image, and text and image; the query feature data calculation module specifically is used for: When the type of the query instruction is text, obtaining the query data text feature data based on the LSTM network; When the type of the query instruction is image, obtaining the text feature data in the query data image and the query image feature data according to the query data, and fusing them to obtain the query data image feature data; When the type of the query instruction is text and image, obtaining the query image data and the query text data according to the query data, and calculating the query image data features and the query text spliced data features to obtain the fused query feature data; The query feature data includes one or more of the query data text feature data, the query data image feature data, and the fused query feature data.

9. An electronic device, characterized in that, It includes: A memory for storing a computer program; A processor for implementing the steps of a cross-modal patent image retrieval method as described in any one of claims 1 to 5 when executing the computer program.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by the processor, the steps of a cross-modal patent image retrieval method as described in any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Cross-modal information retrieval method based on semantic fusion

    CN113536067A

  • Hash retrieval method, system and equipment based on multi-source biological data and medium

    CN116825210A

  • Information retrieval method and device, equipment, program product and storage medium

    CN116975340A

  • System and a method for semantic level image retrieval

    US20200334486A1