A method and system for appearance patent image retrieval based on customizable semantics

By constructing a hybrid mapping design patent retrieval system that combines text and image features, the system addresses the issues of insufficient accuracy and efficiency in existing design patent retrieval technologies, achieving efficient and accurate design patent retrieval results.

CN116415021BActive Publication Date: 2026-02-13SUN YAT SEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211723818.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-30
Publication Date
2026-02-13
Estimated Expiration
2042-12-30

AI Technical Summary

Technical Problem

In existing design patent search technologies, text-based methods require a large amount of data for training to obtain accurate semantic mapping, while image-based methods are inefficient and difficult to customize, resulting in insufficient search accuracy and efficiency.

Method used

A hybrid mapping approach is adopted. By constructing and training an initial screening model and a final screening model for design patents, and combining text and image features, the semantic features of design patent images are generated and mapped using ResNet and Glove word encoders, long short-term memory networks, and other technologies. Generative adversarial neural networks are used to generate hand-drawn sketches to achieve hybrid retrieval of text and images.

Benefits of technology

It improves the accuracy and efficiency of design patent retrieval, reduces the difficulty of image search, enhances the ability to retrieve detailed information, and achieves efficient design patent retrieval results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116415021B_ABST
    Figure CN116415021B_ABST
Patent Text Reader

Abstract

The application discloses a kind of appearance patent image retrieval method and system based on customizable semantics.It includes: from appearance patent database, text-patent pair is filtered and obtained, and the hand-drawn sketch corresponding to the appearance patent image of each patent is generated;ResNet method is used to carry out semantic coding to the appearance patent image, and construct semantic feature library;Appearance patent preliminary screening model is constructed and trained;Appearance patent final screening model is constructed and trained;User inputs the keyword text and hand-drawn sketch to be searched, and keyword text is used as input to carry out preliminary screening, and customizable hand-drawn sketch is used as input to carry out final screening, to obtain accurate appearance patent search result.The application is based on the mode of artificial intelligence, mixed use text and image to the mixed mapping of patent, simultaneously realize the efficiency and accuracy of appearance patent search, through customizable semantic hand-drawn sketch search, more effectively reduce the difficulty of image search, improve the ability of search details.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of image retrieval, image understanding, patent retrieval, and in particular to a method and system for appearance patent image retrieval based on customizable semantics. BACKGROUND

[0002] Intellectual property is the basis of the development of innovative economy, and plays a decisive role in the development of national economic construction. Efficient analysis and use of patent technology will protect and promote the development of industry and enhance the innovation ability of the country and enterprises. Through reasonable analysis of patent technology documents, unnecessary expenses can be avoided and research and development progress can be accelerated. As a creative design patent, appearance design patent is closely related to the development of enterprises. Excellent appearance design will directly affect people's aesthetic habits, promote economic consumption, create new economic growth, and enhance the competitiveness of enterprises. Due to the rapid replacement of products, the number of appearance patents is growing rapidly, and due to the easy-to-mimic nature of appearance design, appearance patents are easily infringed. How to maintain the legality of the entire life cycle of appearance design patents, including patent application, review and rights protection, and how to efficiently manage appearance intellectual property have become an imminent problem. Among them, appearance patent retrieval is an important part of the realization of the popularization and application of intellectual property creation and protection, intelligent review, value assessment and transaction scenarios. How to quickly retrieve interested appearance patents and improve the rationality and legality of the entire patent life cycle has attracted the interest of researchers.

[0003] Appearance patent retrieval provides interested appearance patents for users through various means, providing convenience for patent application, review and rights protection and other processes. The existing main method is to retrieve appearance patents with related words from the appearance patent database according to keywords. However, due to the lack of sufficient text description of appearance patents, it is difficult to obtain ideal retrieval results by retrieving only keywords. Overall, people need a lot of professional knowledge to retrieve related appearance patents, and the retrieval speed is fast, but the retrieval accuracy cannot be guaranteed.

[0004] In recent years, a new round of technological revolution and industrial change is emerging. The formation of big data, the innovation of theoretical algorithm, the improvement of computing power and the evolution of network facilities drive the development of artificial intelligence into a new stage. Intelligence has become an important direction of technology and industrial development. In appearance patent retrieval, intelligence is also an inevitable trend, and a series of patent retrieval methods based on artificial intelligence have emerged. These methods mainly input appearance-text data pairs, use artificial intelligence models for training, and generate semantic mapping of text content to appearance patents. However, due to the development of artificial intelligence-related technologies, such methods require a large amount of manually annotated data to obtain enough learnable data information to complete the training of the model. It is difficult to obtain accurate semantic mapping using a limited data set.

[0005] One of the current prior art is a keyword-based appearance patent retrieval system and method. First, the patent is classified by expert knowledge or patent classification, and then the corresponding appearance patent content is obtained by text comparison. The disadvantage of this method is that the search range is limited to the text content in the patent, and the search method is not flexible enough, and the precision is not good enough.

[0006] The second current prior art is a text semantic-based appearance patent retrieval system and method. First, a text-patent pair is generated by expert knowledge, then a deep neural network is used to learn the mapping relationship from text to patent, and the trained mapping relationship from text to patent is used to search the current user input to obtain the corresponding appearance patent content. The disadvantage of this method is that enough learnable data information is needed to complete the training of the model. If a limited data set is used, it will be difficult to obtain accurate semantic mapping.

[0007] The third current prior art is a method for retrieving appearance patents based on images in the paper "Large scale patent drawing recognition and retrieval". The steps are as follows: first, generate an image-patent pair by expert knowledge, then use a deep neural network to learn the mapping relationship from image to patent, and use the trained mapping relationship to search the current user input image to obtain the corresponding appearance patent content. The disadvantage of this method is that the comparison between images is slow, which greatly reduces the efficiency of the retrieval process. And due to the difficulty of image drawing, lack of customization, makes it difficult to retrieve patents by image. SUMMARY

[0008] The application aims to overcome the shortcomings of the prior art and provides an appearance patent image retrieval method and system based on customizable semantics. The main problem solved by the application is that the existing text-based appearance patent retrieval technology establishes a text-to-patent mapping relationship by training a network model, but requires sufficient learnable data information to complete the training of the model. If a limited data set is used, it will be difficult to obtain a sufficiently accurate semantic mapping. The existing image-based appearance patent retrieval technology establishes an image-to-patent semantic mapping relationship by training a network model, but due to the slow comparison speed between images and the difficulty in drawing images, the lack of customization makes the retrieval difficult and the retrieval efficiency poor. That is, how to break through the limitations of low-efficiency mapping learning from text to patent and from image to patent based on artificial intelligence, and simultaneously realize the efficiency and accuracy of appearance patent retrieval by mixed use of text and image-to-patent hybrid mapping.

[0009] To solve the above problems, the application provides an appearance patent image retrieval method based on customizable semantics, which comprises the following steps:

[0010] Filtering and obtaining text (t)-patent (p) pairs from an appearance patent database, and generating a hand-drawn sketch corresponding to each appearance patent image of the patent (p);

[0011] Using a ResNet method to perform semantic encoding on the appearance patent image, using a pre-trained convolutional neural network ResNet to obtain the semantic features of the appearance patent image p and storing and constructing a semantic feature library;

[0012] Constructing and training an appearance patent primary screening model, first using a Glove word encoder to encode the text to obtain t', and then using a long short-term memory network method to perform semantic encoding on the text to obtain text semantic features Then using a convolutional neural network to map the appearance patent image semantic features and the text semantic features to the same semantic subspace, respectively obtaining and Finally, using a loss function to train the network model and forming an appearance patent primary screening model;

[0013] Constructing and training an appearance patent final screening model, using a ResNet method to obtain the semantic features of the appearance hand-drawn sketch s Then mapping the appearance patent image semantic features and the sketch semantic features to the same semantic subspace, respectively obtaining and Finally, using the obtained As a target vector, make And Can keep semantic labels, and train to form appearance patent final screening model;

[0014] The user inputs the keyword text to be searched and the hand-drawn sketch to be searched. First, the keyword text is taken as the input of the appearance patent primary screening model for searching to initially screen the appearance patent. Then, the customizable semantic hand-drawn sketch is taken as the input of the appearance patent final screening model for searching within the range of the initially screened appearance patent to complete the final screening of the appearance patent, so as to obtain accurate appearance patent search results.

[0015] Preferably, the text (t)-patent (p) pairs are screened and obtained from the appearance patent database, and the hand-drawn sketch corresponding to the appearance patent image of each patent (p) is generated, specifically:

[0016] The appearance patents in the database are classified by using the Locarno classification method, and the text (t)-patent (p) pairs of the classified names and the appearance patents are obtained for subsequent network model training;

[0017] The generated sketch s of the appearance patent image is generated by using a generative adversarial neural network. The method consists of two generation models, two discrimination models and a classifier. The generator in the generative adversarial network uses a u-net structure. One generator uses a real image as input to generate a corresponding hand-drawn sketch. The other generator uses a hand-drawn sketch as input to generate a real image. The discriminator in the generative adversarial network uses a convolutional neural network to judge the generated hand-drawn sketch and image. If the current image is generated by a computer, it is determined as 0, otherwise it is set as 1. The classifier in the generative adversarial network uses a convolutional neural network to classify the generated image.

[0018] Preferably, the appearance patent image is semantically encoded by using a ResNet method, and a pre-trained convolutional neural network ResNet is used to obtain the semantic features of the appearance patent image p and stored, to build a semantic feature library, specifically:

[0019] The method uses a convolutional neural network to train a semantic feature map, classifies and regresses the semantic feature map to obtain the visual semantics in the image, uses a pre-trained convolutional neural network ResNet to obtain the semantic features of the appearance patent image p and store them for subsequent network model training and use, and the specific formula is as follows:

[0020]

[0021] Preferably, the appearance patent primary screening model is constructed and trained, first, the text is encoded by the Glove word encoder to obtain t', and then the text is semantically encoded by the long short-term memory network method to obtain the semantic features of the text The appearance patent image semantic features And the text semantic features are mapped to the same semantic subspace to obtain And Finally, the network model is trained using a loss function, and the appearance patent primary screening model is formed, specifically:

[0022] The text is encoded by the Glove word encoder to obtain t', and then the text is semantically encoded by the long short-term memory network method, which uses a recurrent neural network to train the semantic feature vector, and classifies and trains the semantic feature vector to obtain the semantic information in the text. Specifically, a pre-trained recurrent neural network LSTM is used to obtain the semantic features of the text information t' The specific formula is as follows:

[0023]

[0024] Three loss functions are used to train the network model to ensure that the corresponding functions can be performed. First, a modal classifier is used as a network discriminator to distinguish whether the features are from an image or a text. If it is from an image, the label is assigned as 01; if it is from a text, the label is assigned as 10. It is designed as a 3-layer convolutional network, and the loss function is:

[0025]

[0026] Where m is the true modal label of each sample, and D is a feedforward network composed of fully connected layers;

[0027] Then, in order to ensure that the projected features have discriminative power for semantic labels, the scheme uses a modal intra-discriminative loss function:

[0028]

[0029] Where y is the true modal label of each sample, and M is a feedforward network composed of fully connected layers;

[0030] In order to ensure that the projected features can adapt to the modal difference, that is, to solve the gap problem between modalities, the following loss function is designed. First, generate sample pairs with the same semantic label but different modalities And form a three-tuple sample And Then calculate the multi-modal semantic invariance loss across image and text modalities:

[0031]

[0032]

[0033] Where μ is a control coefficient to prevent data overflow, and l2 represents the L2 distance, used to evaluate the differences between cross-modal features:

[0034]

[0035] M1 and M2 are feedforward networks composed of fully connected layers, and the final intermodal loss function is defined as:

[0036] L imi =L imi,T +L imi,P ;

[0037] To prevent the network model from overfitting, the following loss function is used:

[0038]

[0039] Where l represents the number of layers in the network model, W represents the parameter set of the network, and F represents the norm;

[0040] Ultimately, the entire network is trained using a combination of the three proposed loss functions until the results of the loss functions converge:

[0041] L=αL imi +βL imd +L reg ,

[0042] The hyperparameters α and β control the contributions of the two terms.

[0043] Preferably, a final screening model for design patents is constructed and trained, and the semantic features of the hand-drawn design sketches s are obtained using the ResNet method. Then the semantic features of the appearance patent image and the sketch semantic features Mapped to the same semantic subspace, respectively obtain and Finally, utilize the obtained... As the target vector, and It can retain semantic tags and be trained to form the final screening model for appearance patents, specifically:

[0044] The ResNet method is used to encode the hand-drawn sketch, the method uses a convolutional neural network to obtain a feature map, and the feature map is classified and regressed to obtain the visual semantics in the image, specifically, a convolutional neural network ResNet is used to obtain the semantic features of the appearance hand-drawn sketch s The specific formula is as follows:

[0045]

[0046] Using an attention model can learn different weights for different regions of a given sketch / picture, and the appearance patent image semantic features and the sketch semantic features are mapped to the same semantic subspace, respectively obtaining and In order to better train the network, there are with the same semantic label and with different semantic labels

[0047] In order to ensure that the projected features can adapt to the modal difference, that is, to solve the gap problem between modalities, the following loss function is designed:

[0048]

[0049] In order to enable the network model to distinguish whether the image is an appearance patent image or a hand-drawn sketch, the following loss function is used:

[0050]

[0051] Wherein, l0 and l1 are defined as:

[0052]

[0053] Wherein, R represents a gradient reversal layer, the training objectives before and after the network are opposite, so that the network model can map the appearance patent image or the hand-drawn sketch to the same semantic subspace as much as possible, and the network model can also accurately judge the source of the picture, ψ and φ are a kind of feedforward network composed of full connection layer, i represents sample size, o belongs to 0 or 1, which respectively indicates sketch and real image;

[0054] The obtained is used as a target vector, so that and can retain semantic labels, and the following loss function is used to realize this goal:

[0055]

[0056] wherein, g() represents a semantic reconstruction network, mapping the semantic space into the same semantic space as the text s;

[0057] The entire network is finally trained by the combination of the three proposed loss functions until the results of the loss functions reach convergence:

[0058] L = a1L1 + a2L2 + a3L3,

[0059] wherein the hyperparameters a1, a2 and a3 control the contribution of the three terms.

[0060] Correspondingly, the application also provides a customizable semantic-based appearance patent image retrieval system, comprising:

[0061] a hand-drawn sketch generation unit for screening and obtaining text (t)-patent (p) pairs from an appearance patent database, and generating a hand-drawn sketch corresponding to the appearance patent image of each patent (p);

[0062] a feature library construction unit for performing semantic encoding on the appearance patent image using the ResNet method, using a pre-trained convolutional neural network ResNet to obtain the semantic features of the appearance patent image p and storing, constructing a semantic feature library;

[0063] an appearance patent initial screening model construction and training unit for constructing and training an appearance patent initial screening model, first encoding the text using a Glove word encoder to obtain t', then performing semantic encoding on the text using a long short-term memory network method to obtain text semantic features then using a convolutional neural network to map the semantic features of the appearance patent image and the semantic features of the text to the same semantic subspace, respectively obtaining and Finally, the network model is trained using a loss function, and an appearance patent initial screening model is formed;

[0064] an appearance patent final screening model construction and training unit for constructing and training an appearance patent final screening model, using the ResNet method to obtain the semantic features of the appearance hand-drawn sketch s then mapping the semantic features of the appearance patent image and the semantic features of the sketch to the same semantic subspace, respectively obtaining and Finally, the obtained is used as a target vector, so that and The semantic label can be reserved, and a final appearance patent screening model is trained to form;

[0065] The appearance patent retrieval and result output unit is used for inputting the keyword text to be retrieved and the hand-drawn sketch to be retrieved by the user. Firstly, the keyword text is taken as the input of the appearance patent primary screening model to perform retrieval to initially screen the appearance patent. Then, the customizable semantic hand-drawn sketch is taken as the input of the appearance patent final screening model to perform retrieval within the range of the initially screened appearance patent to complete the final screening of the appearance patent, so that the accurate appearance patent retrieval result is obtained.

[0066] The present application has the following beneficial effects:

[0067] The present application is based on the artificial intelligence mode, is no longer limited to the low-precision mapping learning from text to patent, and is also not limited to the low-efficiency mapping learning from image to patent, but uses the mixed mapping from text and image to patent, so that the efficiency and precision of the appearance patent retrieval are realized. The present application replaces the traditional image search with the hand-drawn sketch search with customizability by virtue of the characteristics that the image such as the hand-drawn sketch is usually easier to convey important scientific or technical information than the text, so that the difficulty of the image search is more effectively reduced, and the ability of the retrieval details is improved. BRIEF DESCRIPTION OF DRAWINGS

[0068] Figure 1 is the flowchart of the appearance patent retrieval method based on the customizable semantic of the embodiment of the present application;

[0069] Figure 2 is the structure diagram of the appearance patent retrieval system based on the customizable semantic of the embodiment of the present application. DETAILED DESCRIPTION

[0070] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.

[0071] Figure 1 is the flowchart of the appearance patent retrieval method based on the customizable semantic of the embodiment of the present application, as Figure 1 shown, the method comprises:

[0072] S1, screening and acquiring the text (t)-patent (p) pair from the appearance patent database, and generating the hand-drawn sketch corresponding to the appearance patent image of each patent (p);

[0073] S2, using ResNet method to perform semantic encoding on the appearance patent image, using a pre-trained convolutional neural network ResNet to obtain the semantic features of the appearance patent image p and store, build a semantic feature library;

[0074] S3, build and train the appearance patent primary screening model, first encode the text using the Glove word encoder to obtain t', and then perform semantic encoding on the text using the long short-term memory network method to obtain the text semantic features Then use the convolutional neural network to map the appearance patent image semantic features and the text semantic features to the same semantic subspace, respectively obtaining and Finally, use the loss function to train the network model and form the appearance patent primary screening model;

[0075] S4, build and train the appearance patent final screening model, use ResNet method to obtain the semantic features of the appearance hand-drawn sketch s Then map the appearance patent image semantic features and the sketch semantic features to the same semantic subspace, respectively obtaining and Finally, use the obtained t' as the target vector, so that and can retain semantic labels, and train to form the appearance patent final screening model;

[0076] S5, the user inputs the keyword text to be searched and the hand-drawn sketch to be searched, first takes the keyword text as the input of the appearance patent primary screening model, performs retrieval to initially screen the appearance patent, then takes the customizable semantic hand-drawn sketch as the input of the appearance patent final screening model, performs retrieval within the range of the initially screened appearance patent to complete the final screening of the appearance patent, thereby obtaining the accurate appearance patent retrieval result.

[0077] Step S1 is as follows:

[0078] S1-1, classifies the appearance patents in the database using the Locarno classification method, obtains the classification name and the text (t)-patent (p) pair of the appearance patent, which is used for subsequent network model training;

[0079] S1-2, a hand-drawn sketch s of an appearance patent image is generated using a generative adversarial neural network, the method consists of two generative models, two discriminative models and a classifier, the generator in the generative adversarial network uses a u-net structure, one generator uses a real image as input to generate a corresponding hand-drawn sketch, and the other generator uses a hand-drawn sketch as input to generate a real image, the discriminator in the generative adversarial network uses a convolutional neural network to judge the generated hand-drawn sketch and image, if the current image is judged to be computer-generated, it is determined as 0, otherwise it is set as 1, and the classifier in the generative adversarial network uses a convolutional neural network to classify the generated image.

[0080] Step S2, specifically as follows:

[0081] S2-1, the method uses a convolutional neural network to obtain semantic feature maps, classifies and regresses the semantic feature maps to obtain visual semantics within the image, and uses a pre-trained convolutional neural network ResNet to obtain semantic features of the appearance patent image p and stores them for subsequent training and use of the network model, the specific formula is as follows:

[0082]

[0083] Step S3, specifically as follows:

[0084] S3-1, the text is encoded using a Glove word encoder to obtain t', and then the text is semantically encoded using a long short-term memory network method, which uses a recurrent neural network to obtain semantic feature vectors, classifies and trains the semantic feature vectors to obtain semantic information within the text, specifically, a pre-trained recurrent neural network LSTM is used to obtain semantic features of the text information t' The specific formula is as follows:

[0085]

[0086] S3-2, three loss functions are used to train the network model to ensure that it can perform corresponding functions, first, a modal classifier is used as a network discriminator to distinguish whether the features come from an image or a text, if from an image, the label is assigned as 01; if from a text, the label is assigned as 10, which is designed as a 3-layer convolutional network, and the loss function is:

[0087]

[0088] Where m is the true modal label of each sample, and D is a feedforward network composed of fully connected layers;

[0089] S3-3, then, in order to ensure that the projected features have discriminability to semantic labels, the scheme uses a modal intra-discriminability loss function:

[0090]

[0091] Where y is the true modal label of each sample, M is a feedforward network composed of fully connected layers;

[0092] S3-4, in order to ensure that the projected features can adapt to modal differences, that is, to solve the problem of the gap between modalities. The following loss function is designed. First, generate sample pairs of the same semantic label and different modalities And form a three-tuple sample And Then calculate the multi-modal semantic invariance loss across image and text modalities:

[0093]

[0094]

[0095] Where μ is a control coefficient to prevent data overflow, and l2 represents the L2 distance used to evaluate the difference between cross-modal semantic features:

[0096]

[0097] Where M1 and M2 are a feedforward network composed of fully connected layers, and the final inter-modal loss function is defined as:

[0098] L imi =L imi,T +L imi,P ;

[0099] S3-5, in order to prevent network model overfitting, the following loss function is used:

[0100]

[0101] Where l represents the number of layers of the network model, W represents the parameter set of the network, and F represents the norm;

[0102] S3-6, the entire network is trained by combining the three proposed loss functions until the result of the loss function converges:

[0103] L=αL imi +βL imd +L reg ,

[0104] Where the hyperparameters α and β control the contribution of the two terms.

[0105] Step S4 is as follows:

[0106] S4-1, as the target vector, makes and The semantic label can be retained, specifically:

[0107] The ResNet method is used to encode the hand-drawn sketch. The method uses a convolutional neural network to obtain a feature map, and performs classification and regression training on the feature map to obtain the visual semantics in the image. Specifically, a convolutional neural network ResNet is used to obtain the semantic features of the appearance hand-drawn sketch s The specific formula is as follows:

[0108]

[0109] S4-2, using an attention model can learn the characteristics of different weights of different regions of a given sketch / picture. The appearance patent image semantic features and the sketch semantic features are mapped to the same semantic subspace, respectively obtaining and In order to better train the network, there are with the same semantic label and

[0110] S4-3, in order to ensure that the projected features can adapt to the modal difference, that is, to solve the gap problem between modalities, the following loss function is designed:

[0111]

[0112] S4-4, in order to enable the network model to distinguish whether the source of the image is an appearance patent image or a hand-drawn sketch, the following loss function is used:

[0113]

[0114] Wherein, l0 and l1 are defined as:

[0115]

[0116] Wherein, R represents the gradient reversal layer, the training objectives before and after the network are opposite, so that the network model can map the appearance patent image or the hand-drawn sketch to the same semantic subspace as much as possible, while the network model can also accurately judge the source of the picture, ψ and φ are a kind of feedforward network composed of full connection layer, i represents sample size, o belongs to 0 or 1, which respectively indicates sketch and real image;

[0117] S4-5, using the obtained As a target vector, make and The semantic label can be retained, and the following loss function is used to realize this goal:

[0118]

[0119] Wherein, g() represents a semantic reconstruction network that maps the semantic space to the same semantic space as the text s;

[0120] S4-6, the entire network is trained by combining the three proposed loss functions until the results of the loss functions converge:

[0121] L = a1L1 + a2L2 + a3L3,

[0122] Where the hyperparameters a1, a2 and a3 control the contribution of the three terms.

[0123] Correspondingly, the application also provides a customizable semantic-based appearance patent image retrieval system, as shown in Figure 2 It comprises:

[0124] A hand-drawn sketch generation unit 1 is used to filter and obtain text (t)-patent (p) pairs from the appearance patent database, and generate a hand-drawn sketch corresponding to the appearance patent image of each patent (p);

[0125] Specifically, the appearance patents in the database are classified using the Locarno classification method, and the text (t)-patent (p) pairs of the classified names and appearance patents are obtained for subsequent network model training; a generative adversarial neural network is used to generate a hand-drawn sketch s of the appearance patent image, the network consists of two generative models, two discriminative models and a classifier, the generators in the generative adversarial network all use the u-net structure, one of which uses real images as input to generate corresponding hand-drawn sketches, the other uses hand-drawn sketches as input to generate real images, the discriminators in the generative adversarial network use convolutional neural networks to judge the generated hand-drawn sketches and images, if the current image is generated by the computer, it is determined as 0, otherwise it is set to 1, and the classifier in the generative adversarial network uses a convolutional neural network to classify the generated images.

[0126] A feature library construction unit 2 is used to use the ResNet method to perform semantic encoding on the appearance patent image, and a pre-trained convolutional neural network ResNet is used to obtain the semantic features of the appearance patent image p and store, and construct a semantic feature library;

[0127] Specifically, the feature map is obtained by training using a convolutional neural network, the feature map is classified and regression trained to obtain visual semantics in the image, and a pre-trained convolutional neural network ResNet is used to obtain semantic features of the appearance patent image p and stored for subsequent training and use of the network model.

[0128] The appearance patent initial screening model construction and training unit 3 is used to construct and train the appearance patent initial screening model. First, the text is encoded using a Glove word encoder to obtain t', and then the text is semantically encoded using a long short-term memory network method to obtain text semantic features The appearance patent image semantic features and the text semantic features are mapped to the same semantic subspace to obtain and Finally, the network model is trained using a loss function, and an appearance patent initial screening model is formed.

[0129] Specifically, the text is encoded using a Glove word encoder to obtain t', and then the text is semantically encoded using a long short-term memory network method. This method uses a recurrent neural network to train a feature vector, classifies the feature vector, and obtains semantic information in the text. Specifically, a pre-trained recurrent neural network LSTM is used to obtain semantic features of the text information t' To distinguish whether the features are from an image or text, if from an image, assign a label of 01, and if from text, assign a label of 10. A 3-layer convolutional network is designed to ensure that the projected features have discriminative power for semantic labels. The projected features can adapt to the differences between modalities, i.e., solve the gap problem between modalities. Three loss functions are used to train the network model to ensure that the corresponding functions can be performed. The loss function is designed to prevent overfitting of the network model. Finally, the entire network is trained by combining the three proposed loss functions until the results of the loss function converge.

[0130] The appearance patent final screening model construction and training unit 4 is used to construct and train the appearance patent final screening model. The ResNet method is used to obtain semantic features of the appearance hand-drawn sketch s The appearance patent image semantic features and the sketch semantic features are mapped to the same semantic subspace to obtain and Finally, the obtained is used as a target vector to make and The semantic label can be retained, and a final appearance patent screening model is trained;

[0131] Specifically, the ResNet method is used to encode the hand-drawn sketch, the feature map is obtained by training the convolutional neural network, and the visual semantics in the image are obtained by classification and regression training of the feature map. The attention model is used to learn different weights of different regions of a given sketch / picture, and the appearance patent image semantic features and the sketch semantic features are mapped to the same semantic subspace, respectively. and In order to better train the network, there are with the same semantic label and with different semantic labels In order to ensure that the projected features can adapt to the modal difference, that is, to solve the gap problem between the modes; the network model can distinguish whether the source of the image is an appearance patent image or a hand-drawn sketch; the obtained t' is used as a target vector, so that and The semantic label can be retained, and three loss functions are used to train the network model respectively to ensure that the corresponding functions can be performed; finally, the entire network is trained by the three proposed loss functions until the result of the loss function converges.

[0132] The appearance patent retrieval and result output unit 5 is used for user to input the keyword text to be retrieved and the hand-drawn sketch to be retrieved. First, the keyword text is used as the input of the appearance patent initial screening model for retrieval to initially screen the appearance patent. Then, the customizable semantic hand-drawn sketch is used as the input of the appearance patent final screening model for retrieval within the range of the initially screened appearance patent to complete the final screening of the appearance patent, so as to obtain accurate appearance patent retrieval results.

[0133] Therefore, based on the artificial intelligence mode, the present application is no longer limited to low-precision mapping learning from text to patent, nor is it limited to low-efficiency mapping learning from image to patent. Instead, the present application uses a hybrid mapping from text and image to patent, while realizing the efficiency and precision of appearance patent retrieval. The present application takes advantage of the fact that images such as hand-drawn sketches can usually convey important scientific or technical information more easily than text, and replaces the traditional image search with a customizable hand-drawn sketch search, which more effectively reduces the difficulty of image search and improves the ability to search details.

[0134] The above describes in detail the appearance patent image retrieval method and system based on customizable semantics provided by the embodiment of the present application, and the principle and implementation manner of the present application are described by using specific examples; the above embodiment description is only used to help understand the method of the present application and its core idea; meanwhile, for the general technical personnel in the art, the specific implementation manner and application range will be changed according to the idea of the present application, and the above description should not be understood as the limitation of the present application.

Claims

1. A method for appearance patent image retrieval based on customizable semantics, characterized in that, The method comprises: Filtering and obtaining text (t)-patent (p) pairs from the appearance patent database, and generating hand-drawn sketches corresponding to the appearance patent images of each patent (p); The appearance patent image is semantically coded by using a ResNet method, and a pre-trained convolutional neural network ResNet is used to obtain semantic features of the appearance patent image p and storage, build a semantic feature library; The appearance patent primary screening model is constructed and trained, first, a text is encoded by using a Glove word encoder to obtain t', and then a long short-term memory network method is used for semantic encoding of the text to obtain a text semantic feature The appearance patent image semantic feature and the text semantic feature are mapped to the same semantic subspace by using a convolutional neural network, and and finally, a loss function is used for training the network model, and the appearance patent primary screening model is formed. The appearance patent final screening model is constructed and trained, and a ResNet method is used to obtain semantic features of the appearance hand-drawn sketch s Then the appearance patent image semantic features and the sketch semantic features are mapped to the same semantic subspace, respectively obtaining and Finally, the obtained is used as a target vector, so that and can retain semantic labels, and the appearance patent final screening model is trained. The user inputs the keyword text to be searched and the hand-drawn sketch to be searched. First, the keyword text is used as the input of the appearance patent initial screening model to perform retrieval to initially screen the appearance patent. Then, the customizable semantic hand-drawn sketch is used as the input of the appearance patent final screening model to perform retrieval within the range of the initially screened appearance patent to complete the final screening of the appearance patent, so as to obtain accurate appearance patent search results.

2. The customizable semantic-based appearance patent image retrieval method of claim 1, wherein, The filtering and obtaining text (t)-patent (p) pairs from the appearance patent database, and generating hand-drawn sketches corresponding to the appearance patent images of each patent (p) are specifically: Classify the appearance patents in the database using the Locarno classification method to obtain the classification name and the text (t)-patent (p) pairs of the appearance patent for subsequent network model training; Generate the hand-drawn sketch s of the appearance patent image using the generative adversarial neural network. The method consists of two generative models, two discriminative models, and a classifier. The generator in the generative adversarial network uses the u-net structure. One generator uses a real image as input to generate a corresponding hand-drawn sketch. The other generator uses a hand-drawn sketch as input to generate a real image. The discriminator in the generative adversarial network uses a convolutional neural network to judge the generated hand-drawn sketch and image. If the current image is generated by a computer, it is determined as 0, otherwise it is set as 1. The classifier in the generative adversarial network uses a convolutional neural network to classify the generated image.

3. The customizable semantic-based appearance patent image retrieval method of claim 1, wherein, The ResNet method is used for semantic coding of the appearance patent image, and a pre-trained convolutional neural network ResNet is used to obtain semantic features of the appearance patent image p And store, build semantic feature library, specifically: The method utilizes a convolutional neural network to obtain a semantic feature map, classifies and regresses the semantic feature map, obtains visual semantics in the image, and uses a pre-trained convolutional neural network ResNet to obtain semantic features of the appearance patent image p and stores them for subsequent training and use of the network model, and the specific formula is as follows:

4. The customizable semantic-based appearance patent image retrieval method of claim 1, wherein, The appearance patent primary screening model is constructed and trained, first, a Glove word encoder is used to encode the text to obtain t', and then a long short-term memory network method is used to encode the text semantics to obtain text semantic features The appearance patent image semantic features and the text semantic features are mapped to the same semantic subspace to obtain and Finally, the network model is trained using a loss function, and the appearance patent primary screening model is formed, specifically: The text is encoded by using the Glove word encoder to obtain t', and then the long short-term memory network method is used for semantic encoding, the method uses a recurrent neural network to train to obtain a semantic feature vector, and the semantic feature vector is classified and trained to obtain semantic information in the text. The specific formula is as follows: Three kinds of loss functions are used to train the network model to ensure that it can perform corresponding functions. First, a modal classifier is used as the discriminator of the network to distinguish whether the feature is from an image or text. If it is from an image, the label is assigned as 01; if it is from text, the label is assigned as 10. It is designed as a 3-layer convolutional network, and the loss function is: Where m is the true modal label of each sample, and D is a feedforward network consisting of fully connected layers; Then, in order to ensure that the projected features have discriminative power for semantic labels, a modal intra-discriminative loss function is used: Where y is the true modal label of each sample, and M is a feedforward network consisting of fully connected layers; To ensure that the projected features can adapt to the modality difference, i.e., solve the gap problem between modalities, a loss function is designed, specifically, first, generate sample pairs of the same semantics but different modalities and form a triple sample and Then, the multi-modal semantic invariance loss across image and text modalities is calculated: Where μ is a control coefficient to prevent data overflow, and l2 represents the L2 distance for evaluating the difference between cross-modal semantic features: Where M1 and M2 are a feedforward network consisting of fully connected layers, and the final inter-modal loss function is defined as: L imi = L imi,T + L imi,P ; In order to prevent the network model from overfitting, the following loss function is used: Where l represents the number of layers of the network model, W represents the parameter set of the network, and F represents the norm; The entire network is trained by combining the three proposed loss functions until the result of the loss function converges: L = aL + βL imi + L imd + L reg , Where the hyperparameters α and β control the contribution of the two terms.

5. The customizable semantic-based appearance patent image retrieval method of claim 1, wherein, The appearance patent final screening model is constructed and trained, and a ResNet method is used to obtain semantic features of the appearance hand-drawn sketch s Then, the appearance patent image semantic features and the sketch semantic features are mapped to the same semantic subspace, respectively obtaining and Finally, the obtained is used as a target vector, so that and can retain semantic labels, and the appearance patent final screening model is trained, specifically: The ResNet method is used for encoding the hand-drawn sketch, the method uses a convolutional neural network to obtain a feature map, and the feature map is classified and regressed to obtain visual semantics in the image, specifically, a convolutional neural network ResNet is used to obtain semantic features of the appearance hand-drawn sketch s The specific formula is as follows: Using attention model can learn different weights of different areas of a given sketch / picture, and map the appearance patent image semantic features and the sketch semantic features to the same semantic subspace, respectively obtaining and In order to better train the network, there are with the same semantic label and different semantic labels In order to ensure that the projected features can adapt to the modal difference, i.e., to solve the gap problem between modalities, the following loss function is designed: In order to enable the network model to distinguish whether the source of the image is an appearance patent image or a hand-drawn sketch, the following loss functions are used: Wherein, l0 and l1 are defined as: Wherein, R represents a gradient reversal layer, the training objectives before and after the network are opposite, so that the network model can map the appearance patent image or the hand-drawn sketch to the same semantic subspace as much as possible, and the network model can also accurately judge the source of the picture, ψ and φ are a kind of feedforward network composed of full connection layers, i represents sample size, o belongs to 0 or 1, which respectively indicates sketch and real image; Using the acquired As a target vector, make so that And It is possible to preserve semantic labels and to manifest this goal using the following loss function: wherein, g() represents a semantic reconstruction network that maps the semantic space into the same semantic space as the text s. Finally, the entire network is trained by combining the three proposed loss functions until the results of the loss functions converge: L = α1L1 + α2L2 + α3L3, Wherein, the hyperparameters α1, α2 and α3 control the contribution of the three terms.

6. A customizable semantic based appearance patent image retrieval system, characterized by, The system comprises: A hand-drawn sketch generation unit for screening and obtaining text (t)-patent (p) pairs from an appearance patent database, and generating a hand-drawn sketch corresponding to each appearance patent image of the patent (p); The feature library construction unit is configured to perform semantic coding on the appearance patent image by using a ResNet method, and a pre-trained convolutional neural network ResNet is used to obtain semantic features of the appearance patent image p and store, and construct a semantic feature library; The appearance patent initial screening model construction and training unit is used for constructing and training the appearance patent initial screening model, first uses a Glove word encoder to encode the text to obtain t', and then uses a long short-term memory network method to encode the text semantics to obtain text semantic features The appearance patent image semantic features and the text semantic features are mapped to the same semantic subspace, respectively obtaining and Finally, the network model is trained using a loss function, and the appearance patent initial screening model is formed. The appearance patent final screening model construction and training unit is configured to construct and train the appearance patent final screening model, and obtain semantic features of the appearance hand-drawn sketch s by using a ResNet method The appearance patent image semantic features and the sketch semantic features are then mapped to the same semantic subspace, respectively obtaining and Finally, the obtained is used as a target vector, so that and can retain semantic labels and be trained to form the appearance patent final screening model; An appearance patent retrieval and result output unit for user input of a keyword text to be retrieved and a hand-drawn sketch to be retrieved, first taking the keyword text as the input of the appearance patent initial screening model to perform retrieval to initially screen the appearance patents, and then taking the customizable semantic hand-drawn sketch as the input of the appearance patent final screening model to perform retrieval within the range of the initially screened appearance patents to complete the final screening of the appearance patents, thereby obtaining accurate appearance patent retrieval results.

7. The customizable semantics based appearance patent image retrieval system as claimed in claim 6, wherein, The hand-drawn sketch generation unit needs to classify the appearance patents in the database using the Locarno classification method, obtain the classification name and the text (t)-patent (p) pair of the appearance patent, and use it for subsequent network model training; generate a hand-drawn sketch s of the appearance patent image using a generative adversarial neural network, the network comprises two generative models, two discriminative models and a classifier, the generators in the generative adversarial network all use the u-net structure, one of the generators uses a real image as input to generate a corresponding hand-drawn sketch, the other generator uses a hand-drawn sketch as input to generate a real image, the discriminators in the generative adversarial network use convolutional neural networks to judge the generated hand-drawn sketch and image, if the current image is judged to be computer generated, it is determined as 0, otherwise it is set as 1, and the classifier in the generative adversarial network uses a convolutional neural network to classify the generated image.

8. The customizable semantics based appearance patent image retrieval system as claimed in claim 6, wherein, The semantic feature library construction unit needs to use a convolutional neural network to obtain a semantic feature map, classify and regress the semantic feature map, obtain visual semantics in the image, and use a pre-trained convolutional neural network ResNet to obtain semantic features of the appearance patent image p And store it for subsequent network model training and use.

9. The customizable semantics based appearance patent image retrieval system as claimed in claim 6, wherein, The appearance patent initial screening model construction and training unit needs to encode the text by using a Glove word encoder to obtain t', and then encode the semantics of the text by using a long short-term memory network method. The method uses a recurrent neural network to train to obtain a feature vector, classifies and trains the feature vector, obtains semantic information in the text, and specifically uses a pre-trained recurrent neural network LSTM to obtain semantic features of the text information t' In order to distinguish whether the feature is from an image or text, if it is from an image, a label of 01 is assigned, and if it is from text, a label of 10 is assigned. A 3-layer convolutional network is designed. The projected features have discriminability for semantic labels. The projected features can adapt to the modal difference, that is, solve the gap problem between modes. Three kinds of loss functions are used to train the network model to ensure that the corresponding functions can be performed. The loss function is designed to prevent the network model from overfitting. Finally, the entire network is trained by combining the three proposed loss functions until the result of the loss function converges.

10. The customizable semantics based appearance patent image retrieval system as claimed in claim 6, wherein, The appearance patent final screening model construction and training unit needs to use the ResNet method to perform semantic coding on the hand-drawn sketch. The method uses a convolutional neural network to obtain a semantic feature map, performs classification and regression training on the semantic feature map, and obtains visual semantics in the image. Specifically, a convolutional neural network ResNet is used to obtain the semantic features of the appearance hand-drawn sketch s Using an attention model can learn different weights for different regions of a given sketch / picture, and the appearance patent image semantic features and the sketch semantic features are mapped to the same semantic subspace to obtain and In order to better train the network, there are with the same semantic label and different semantic labels In order to ensure that the projected features can adapt to the modal difference, that is, to solve the gap problem between the modes; make the network model able to distinguish whether the source of the image is an appearance patent image or a hand-drawn sketch; use the obtained as a target vector, so that and can retain semantic labels, and three kinds of loss functions are used to train the network model respectively to ensure that the corresponding functions can be performed; finally, the entire network is trained by combining the three proposed loss functions until the result of the loss function converges.

Citation Information

Patent Citations

  • Image-text joint representation search method, system, server and storage medium

    CN112883218A

  • System for interfacing between an operator and a virtual object for computer aided design applications

    US20060119578A1