Application suggestion method and system combined with appearance patent image retrieval

By combining the application proposal methods and systems for appearance patent graphic search, it solves the problem that applicants find it difficult to understand the existing appearance design and preparation in an irregular manner, and realizes intelligent and personalized application proposals, improving the quality and efficiency of application materials.

CN118522028BActive Publication Date: 2025-05-16ZHEJIANG ZHIYIBEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410727970.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-06
Publication Date
2025-05-16
Estimated Expiration
2044-06-06

AI Technical Summary

Technical Problem

During the application process of design patents, it is difficult for applicants to fully understand the existing design, lack the grasp of the design trends, the application materials are not standardized, and professional knowledge is lacking, it is difficult to conduct comprehensive analysis and evaluation, and there is a lack of targeted improvement suggestions.

Method used

Provide an application proposal method and system combining appearance patent graphics retrieval, by preprocessing the multi-view diagram of the appearance patent application, extracting the fusion feature vector, searching the closest existing patent, and generating an appearance patent application proposal report.

Benefits of technology

It realizes intelligent and personalized appearance patent graphics application suggestions, can be automatically retrieved, provides targeted application materials preparation guidance, and provides professional analysis, evaluation and improvement suggestions, improving the quality and efficiency of application materials.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118522028B_ABST
    Figure CN118522028B_ABST
Patent Text Reader

Abstract

The present invention provides an application suggestion method and system combined with appearance patent graphic retrieval, which relates to the field of artificial intelligence technology, including preprocessing multi-view graphics of appearance patents to generate unified multi-view graphics, extracting fusion features through a pre-trained multi-view graphic feature extraction network, and determining a fusion feature vector to be applied for; inputting the fusion feature vector to be applied for into an appearance patent graphic search network, searching for the closest existing patent, performing multi-view graphic feature extraction on the graphics in the existing appearance patents, and generating an existing patent feature vector; based on a learned graphic similarity measurement network, calculating the similarity score between the feature vector to be applied for and each existing patent feature vector, determining a reference graphic set according to a preset selection number, and generating an appearance patent application suggestion report through a natural language generation model based on the feature vector to be applied for and the existing patent feature vector corresponding to the reference graphic set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an application suggestion method and system combined with appearance patent graphic retrieval. Background Art

[0002] In the application process of design patents, applicants need to submit design drawings and related documents for examination by the patent examination agency. However, due to the subjectivity and complexity of judging the innovation and patentability of designs, applicants often face many challenges when preparing application materials. They find it difficult to fully understand existing designs and lack a grasp of design trends. Application materials are not prepared in a standardized manner. At the same time, applicants may lack professional knowledge in the field of patents, making it difficult to conduct a comprehensive analysis and evaluation of their designs and lack targeted improvement suggestions.

[0003] In order to help applicants better prepare design patent application materials, there have been some attempts in the prior art, such as providing online search tools to help applicants find relevant existing designs; providing application templates and examples to guide applicants to prepare standardized application documents; providing online consulting and evaluation services, and professionals analyzing and suggesting applicants' design plans. However, there are still the following deficiencies: the degree of intelligence is not high, manual screening and comparison are still required, lack of pertinence and inability to provide personalized guidance and suggestions, and the service quality is affected by manpower and time constraints.

[0004] In summary, there is an urgent need for an intelligent and personalized appearance patent graphic application suggestion method and system, which can automatically search according to the applicant's specific design, provide targeted application materials preparation guidance, and give professional analysis, evaluation and improvement suggestions. The present invention can solve the problems in the prior art. Summary of the invention

[0005] The embodiment of the present invention provides an application suggestion method and system combined with appearance patent graphic retrieval, which can solve the problems in the prior art.

[0006] According to a first aspect of the embodiments of the present invention,

[0007] Provide an application suggestion method combined with appearance patent image search, including:

[0008] Preprocess the multi-view graphics to be applied for the appearance patent to generate a unified multi-view graphics, extract fusion features based on the unified multi-view graphics through a pre-trained multi-view graphics feature extraction network, and determine the fusion feature vector to be applied for;

[0009] Input the fused feature vector to be applied into the appearance patent graphic search network, search for the closest existing patent, extract multi-view graphic features of the graphics in the existing appearance patent, and generate the existing patent feature vector, wherein the appearance patent graphic search network extracts the fused feature vector, performs clustering division, and constructs a graphic search tree;

[0010] Based on the learned graphic similarity measurement network, the similarity score between the feature vector to be applied for and each of the feature vectors of the existing patent is calculated, and a reference graphic set is determined according to a preset selection number. Based on the feature vector to be applied for and the feature vectors of the existing patent corresponding to the reference graphic set, a natural language generation model is used to generate a report on the recommendation for an appearance patent application.

[0011] In an optional embodiment,

[0012] Preprocessing the multi-view graphics to be applied for the appearance patent to generate a unified multi-view graphics, extracting fusion features based on the unified multi-view graphics through a pre-trained multi-view graphics feature extraction network, and determining the fusion feature vector to be applied for includes:

[0013] Based on the semantic segmentation method, a binary mask corresponding to the multi-view graphics is generated, the multi-view graphics are segmented, and the target object graphics are extracted; through key point matching, the rotation, translation and scaling of the target object graphics are eliminated, and the target object alignment graphics are determined, and the target object alignment graphics are scaled according to a preset graphic size to generate a unified multi-view graphics;

[0014] A scale pyramid is constructed based on the unified multi-view graph, each scale graph in the scale pyramid is operated, a local feature corresponding to the scale graph is extracted by a convolution layer to obtain a local feature graph, a transformation layer performs spatial transformation on the local feature graph, a pooling layer performs downsampling on the transformed local feature graph, and a fully connected layer flattens the pooled local feature graph and performs feature transformation and dimension reduction to obtain a multi-scale feature vector;

[0015] The multi-scale feature vector is input into the feature fusion layer, and the scale weight corresponding to each scale feature is learned iteratively, the attention to each scale feature is adaptively adjusted, and a fused feature vector is generated through element-level weighting.

[0016] In an optional embodiment,

[0017] The appearance patent graph search network extracts fusion feature vectors, performs clustering division, and constructs a graph search tree, including:

[0018] Based on the existing appearance patent, a multi-view graphic feature extraction network is applied to extract the fused feature vector of the corresponding graphic, and the existing appearance patent, the graphic and the corresponding fused feature vector are stored based on the vector database;

[0019] Based on the two fused feature vectors, the cosine similarity is calculated, the weight is determined, and a similarity matrix is ​​constructed. The fused feature vectors are used as graph nodes and the weights are used as graph edges to construct an undirected weighted graph and start clustering:

[0020] Traversing each graph node in the undirected weighted graph, determining the sum of weights of all graph edges connected to each graph node, constructing a degree matrix, and calculating a Laplace matrix based on the similarity matrix and the degree matrix;

[0021] Performing eigenvalue decomposition on the Laplace matrix to obtain eigenvalues ​​and corresponding eigenvectors, arranging the eigenvalues ​​from small to large, selecting eigenvalues ​​according to a preset feature selection threshold, determining corresponding eigenvectors, constructing a feature matrix, normalizing the feature matrix by rows, and performing clustering based on the K-Means algorithm to obtain a clustering division result;

[0022] The undirected weighted graph is used as the root node of the tree, and the clustering result is the first-layer tree node. Based on each of the first-layer tree nodes, clustering is iteratively performed to obtain the corresponding clustering result, which is used as the subtree node of the first-layer tree node and the second-layer tree node. The subtree nodes are constructed in a loop in sequence until the preset depth upper limit is met to obtain a multi-level index tree. The centrality of each tree node in the index tree is cyclically calculated to determine the center of the tree node.

[0023] The multi-level index tree, combined with the tree node centers, determines a graph search tree.

[0024] In an optional embodiment,

[0025] Based on the learned graph similarity measurement network, the similarity score between the feature vector to be applied for and each feature vector of the existing patent is calculated, and according to the preset selection number, the reference graph set is determined to include:

[0026] Based on the existing appearance patents, a graphic training data set is determined, and a triplet including an anchor sample, a positive sample, and a negative sample is randomly sampled from the graphic training data set, wherein the positive sample and the anchor sample belong to the same tree node in the appearance patent graphic search tree, and the negative sample and the anchor sample do not belong to the same tree node in the appearance patent graphic search tree;

[0027] Based on the multi-view graphic feature extraction network, extracting anchor point feature vectors corresponding to anchor point samples, positive feature vectors corresponding to positive samples, and negative feature vectors corresponding to negative samples respectively;

[0028] Based on the anchor feature vector and the positive feature vector, determine the positive distance between the anchor sample and the positive sample, calculate the positive sample pair loss, based on the anchor feature vector and the negative feature vector, determine the negative distance between the anchor sample and the negative sample, calculate the negative sample pair loss; based on the positive distance and the negative distance, determine the triplet loss function, based on the positive sample pair loss and the negative sample pair loss, determine the contrast loss function; based on the triplet loss function and the contrast loss function, determine the comprehensive loss function;

[0029] Based on the comprehensive loss function, the gradient descent algorithm is used to iteratively optimize the network parameters of the graph similarity measurement network. The network parameters are updated through back propagation to minimize the comprehensive loss until a preset number of iterations is reached to determine the optimal graph similarity measurement network.

[0030] In an optional embodiment,

[0031] The comprehensive loss function has the following formula:

[0032] ;

[0033] in, a represents the anchor point sample, p represents a positive sample, n represents negative samples, L tr represents the triplet loss function, f ( a ) represents the anchor feature vector, f ( p ) represents a positive eigenvector, f ( n ) represents a negative eigenvector, d ( f ( a ), f ( p )) indicates positive distance, d ( f ( a ), f ( n )) represents a negative distance, magin represents the preset minimum distance hyperparameter, L co represents the contrast loss function, L cb represents the comprehensive loss function, α represents the weight coefficient of the triplet loss function, β Represents the weight coefficient of the contrast loss function.

[0034] In an optional embodiment,

[0035] Randomly sampling from the graphic training data set to construct a triplet containing an anchor sample, a positive sample, and a negative sample includes:

[0036] In each training batch, for each anchor sample, calculate the distance between it and all other samples in the batch. Based on the preset positive sample distance threshold, select the sample with the largest distance to the anchor sample as the candidate difficult positive sample. Based on the preset negative sample distance threshold, select the sample with the smallest distance to the anchor sample as the candidate difficult negative sample.

[0037] In multiple training batches, the frequency of positive samples selected as candidate difficult positive samples and the frequency of negative samples selected as candidate difficult negative samples are counted in turn, and the samples are sorted according to the frequency, and the samples with the highest frequency are selected as the final difficult positive samples and the final difficult negative samples;

[0038] Based on the training results of multiple training batches, the positive sample distance threshold and the negative sample distance threshold are dynamically adjusted to construct a triplet containing anchor samples, positive samples and negative samples.

[0039] According to a second aspect of the embodiments of the present invention,

[0040] Provided is an application suggestion system combined with appearance patent image retrieval, comprising:

[0041] The first unit is used to pre-process the multi-view graphics to be applied for the appearance patent, generate a unified multi-view graphics, extract fusion features based on the unified multi-view graphics through a pre-trained multi-view graphics feature extraction network, and determine the fusion feature vector to be applied for;

[0042] The second unit is used to input the fused feature vector to be applied into the appearance patent graphic search network, search for the closest existing patent, perform multi-view graphic feature extraction on the graphics in the existing appearance patent, and generate the existing patent feature vector, wherein the appearance patent graphic search network extracts the fused feature vector, performs clustering division, and constructs a graphic search tree;

[0043] The third unit is used to calculate the similarity score between the feature vector to be applied for and each of the feature vectors of the existing patent based on the learned graphic similarity measurement network, determine the reference graphic set according to a preset selection number, and generate a design patent application recommendation report based on the feature vector to be applied for and the feature vectors of the existing patents corresponding to the reference graphic set through a natural language generation model.

[0044] According to a third aspect of the embodiments of the present invention,

[0045] An electronic device is provided, comprising:

[0046] processor;

[0047] a memory for storing processor-executable instructions;

[0048] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0049] According to a fourth aspect of the embodiments of the present invention,

[0050] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.

[0051] In the embodiment of the present invention, feature information from different perspectives is integrated to ensure that the extracted feature vector is more comprehensive and accurate, and can better describe the true form of the appearance patent to be applied for; by training the multi-perspective graphic feature extraction network, the model can learn more diverse graphic features, improve its generalization ability, and adapt to more diverse patent application scenarios; by clustering and partitioning to build a graphic search tree, the organization and indexing efficiency of large-scale patent graphic data is improved, making the search process more efficient and accurate; by clustering and partitioning the appearance patent graphics, a graphic search tree is built to effectively organize and index large-scale patent graphic data and improve search efficiency; the structure of the graphic search tree allows searching at multiple levels, narrowing the search scope more quickly, locating the most similar existing patents, reducing search time, and helping to optimize the organization of the feature space, so that similar patent graphics are concentrated together, and further The accuracy and efficiency of retrieval are improved step by step; according to the calculated similarity score, the reference graphic set is determined according to the preset selection number, and the existing patent graphics that are most similar to the patent to be applied for are effectively screened out to ensure the relevance and representativeness of the reference graphics; the anchor feature vector, positive feature vector and negative feature vector contain rich graphic features, which enhances the accuracy of similarity measurement; by calculating the distance between the anchor sample and the positive sample and the negative sample, the positive distance and negative distance are obtained respectively, which can accurately measure the similarity and difference between samples; based on the positive distance and negative distance calculation loss, the network's prediction deviation of the similarity of positive sample pairs and negative sample pairs is measured, which helps to guide network optimization; the triple loss function and the contrast loss function are weightedly combined, and the constraints of the two loss functions are comprehensively considered, so that the network learns a more comprehensive similarity measurement model; by adjusting the weight of the loss function, it can be balanced. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 A flowchart of a method for applying for patent image retrieval combined with an embodiment of the present invention;

[0053] Figure 2 It is a schematic diagram of the structure of the application suggestion system combined with the appearance patent image retrieval according to an embodiment of the present invention. DETAILED DESCRIPTION

[0054] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0055] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0056] Figure 1 FIG. 1 is a flow chart of a method for applying for patent image retrieval in combination with an embodiment of the present invention. Figure 1 As shown, the method includes:

[0057] S101. Preprocessing the multi-view graphics to be applied for the appearance patent to generate a unified multi-view graphics, based on the unified multi-view graphics, extracting fusion features through a pre-trained multi-view graphics feature extraction network, and determining the fusion feature vector to be applied for;

[0058] In this embodiment, by uniformly processing multi-view graphics, the all-round features of the object can be captured, avoiding the problem of inaccurate feature extraction caused by insufficient single-view graphics information; the feature information from different viewpoints is integrated to ensure that the extracted feature vector is more comprehensive and accurate, and can better describe the true form of the appearance patent to be applied for; by training the multi-view graphic feature extraction network, the model can learn more diverse graphic features, improve its generalization ability, and adapt to more diverse patent application scenarios; the introduction of multi-view data reduces the risk of model overfitting, because the model needs to process more diverse data instead of relying solely on information from a single viewpoint.

[0059] In an optional embodiment, the multi-view graphics to be applied for the appearance patent are preprocessed to generate a unified multi-view graphics, and based on the unified multi-view graphics, a pre-trained multi-view graphics feature extraction network is used to extract fusion features, and determining the fusion feature vector to be applied for includes:

[0060] Based on the semantic segmentation method, a binary mask corresponding to the multi-view graphics is generated, the multi-view graphics are segmented, and the target object graphics are extracted; through key point matching, the rotation, translation and scaling of the target object graphics are eliminated, and the target object alignment graphics are determined, and the target object alignment graphics are scaled according to a preset graphic size to generate a unified multi-view graphics;

[0061] A scale pyramid is constructed based on the unified multi-view graph, each scale graph in the scale pyramid is operated, a local feature corresponding to the scale graph is extracted by a convolution layer to obtain a local feature graph, a transformation layer performs spatial transformation on the local feature graph, a pooling layer performs downsampling on the transformed local feature graph, and a fully connected layer flattens the pooled local feature graph and performs feature transformation and dimension reduction to obtain a multi-scale feature vector;

[0062] The multi-scale feature vector is input into the feature fusion layer, and the scale weight corresponding to each scale feature is learned iteratively, the attention to each scale feature is adaptively adjusted, and a fused feature vector is generated through element-level weighting.

[0063] The semantic segmentation method specifically refers to a computer vision technology that is used to classify each pixel in an image into a specific category. Semantic segmentation is different from target detection, which only marks the bounding box of the target. Instead, it is accurate to the pixel level and provides a detailed outline of the target object. When generating a binary mask corresponding to a multi-view graphic, the semantic segmentation method can distinguish between the target object and the background and extract the precise outline of the target object.

[0064] The binary mask specifically refers to an image processing technology that divides pixels in an image into two categories: foreground and background. Usually, a black and white image is used for representation, where foreground pixels correspond to the target object, are represented by white, and have a value of 1, and background pixels are represented by black, and have a value of 0. By generating a binary mask through semantic segmentation, the target object in the image can be separated from the background, providing a basis for subsequent graphic segmentation and target object graphic extraction.

[0065] The key points specifically refer to points in an image that are unique and stable, and are usually used for image matching and alignment. Key points can be corner points, edge points, or other significant feature points in an image. When eliminating the rotation, translation, and scaling of the target object graphics, the corresponding relationship between the target object graphics can be found through key point matching, thereby performing an alignment operation.

[0066] The scale pyramid specifically refers to a multi-scale image analysis method that processes images at different scales to capture features at different scales. Images at multiple scales are usually generated by downsampling the image. Constructing a scale pyramid allows the model to extract features at different scales, thereby obtaining a richer and more robust feature representation.

[0067] The spatial transformation specifically refers to the transformation of the geometric structure of the image during the image processing, such as rotation, translation, scaling, etc. Through the transformation layer, the feature map can be dynamically spatially transformed. Spatial transformation of local feature maps can standardize the geometric structure of the feature map, eliminate the effects of rotation, translation and scaling, and thus improve the accuracy of feature extraction.

[0068] The element-level weighting specifically refers to assigning different weights to each element of the input feature vector, and adjusting the contribution of each element to the final feature vector through these weights. The weights are usually obtained through model learning. In the feature fusion layer, the element-level weighting iteratively learns the weights of each scale feature, adaptively adjusts the attention to each scale feature, so that the generated fused feature vector can more effectively represent the characteristics of the target object.

[0069] First, the input multi-view graphics are processed using semantic segmentation technology. Through the preset semantic segmentation model, a semantic label can be assigned to each pixel to generate a binary mask of the same size as the original graphics. Each pixel value in the binary mask indicates whether the position belongs to the target object. Using the generated binary mask, the target object graphics of interest can be segmented from the original multi-view graphics, removing the interference of background and irrelevant areas.

[0070] Next, in order to eliminate the differences in rotation, translation and scaling of the target object graphics under different viewing angles, graphics alignment is required. By detecting key points in the target object graphics, such as corner points, edge points, etc., and finding the matching relationship between the corresponding key points between graphics of different viewing angles, the parameters of the viewing angle transformation can be estimated. Using the parameters, the target object graphics can be rotated, translated and scaled so that it is aligned to a unified coordinate space under different viewing angles.

[0071] The aligned target object graphics may have scale differences. In order to facilitate subsequent feature extraction, scale normalization is required. According to the preset unified graphic size, the target object alignment graphics are scaled so that the width and height match the preset size, and a multi-view graphics with unified scale is obtained, providing consistent input for feature extraction.

[0072] After obtaining a unified multi-view image, in order to extract local features at different scales, a scale pyramid needs to be constructed. By downsampling the original image multiple times, a series of images with different resolutions can be obtained to form a scale pyramid. Each layer in the pyramid corresponds to a different scale. From the bottom to the top, the resolution of the image gradually decreases, but the receptive field gradually increases.

[0073] For each scale map in the scale pyramid, a convolutional neural network is used to extract features. The convolution layer performs local perception of the image through a sliding window, extracts the features of the local area, and obtains a series of local feature maps, which retain the local pattern and structural information of the image at different positions and scales.

[0074] In order to enhance the expressive power of local features, the local feature map can be spatially transformed. The transformation layer learns a set of affine transformation parameters to perform transformation operations such as rotation, translation, and scaling on the local feature map, making it invariant to spatial changes, which helps to extract more robust and discriminative features.

[0075] The local feature map processed by the transformation layer is downsampled by the pooling layer. The pooling operation aggregates the local area and extracts the maximum or average value in the area, which can reduce the size of the feature map while retaining important feature information, helping to reduce computational complexity and improve the robustness of the feature.

[0076] The pooled local feature map is further transformed and dimensionally reduced through the fully connected layer. The fully connected layer flattens the local feature map into a one-dimensional vector and maps the high-dimensional features into a low-dimensional space by learning a set of weight matrices, removing redundant information and extracting a more compact and discriminative feature representation.

[0077] By performing the above steps on each scale map in the scale pyramid, a series of local feature vectors of different scales can be obtained, and the feature vectors can be fused. The feature fusion layer adaptively adjusts the degree of attention to features of different scales by learning a set of scale weights. The larger the weight, the greater the contribution of the features of the corresponding scale to the final representation. By weighted summing of multi-scale feature vectors, a fused feature vector that comprehensively considers information of different scales can be obtained.

[0078] In this embodiment, the binary mask generated by the semantic segmentation technology can accurately segment the target object graphics, remove the background and irrelevant areas, and ensure that the extracted target object graphics are purer; the method based on key point matching can accurately estimate the perspective transformation parameters to ensure the accuracy of graphics alignment; constructing a scale pyramid can capture the features of the image at different scales, thereby obtaining a richer feature representation; the convolution layer can effectively extract the local features of the image and retain the local pattern and structural information of the image at different positions and scales; the feature fusion layer adaptively adjusts the attention to the features of different scales by learning the weights of the features of each scale, so that the fused feature vector comprehensively considers multi-scale information; the fused feature vector generated by element-level weighting, the scale with a larger weight contributes more to the final representation, thereby enhancing the discriminative ability and robustness of the feature representation.

[0079] S102. Input the fused feature vector to be applied into the appearance patent graphic search network, search for the closest existing patent, extract multi-view graphic features of the graphics in the existing appearance patent, and generate an existing patent feature vector, wherein the appearance patent graphic search network extracts the fused feature vector, performs clustering, and constructs a graphic search tree;

[0080] In this embodiment, by fusing feature vectors into the search network, the existing patents most similar to the patent to be applied for can be quickly and accurately retrieved; multi-view graphic feature extraction and feature vector generation ensure a comprehensive and accurate description of the existing patent graphics, providing a reliable basis for similarity assessment; by clustering and partitioning, a graphic search tree is constructed to improve the organization and indexing efficiency of large-scale patent graphic data, making the search process more efficient and accurate; by clustering and partitioning the appearance patent graphics, a graphic search tree is constructed to effectively organize and index large-scale patent graphic data and improve search efficiency; the structure of the graphic search tree allows searching at multiple levels, narrowing the search scope more quickly, locating the most similar existing patents, reducing search time, and helping to optimize the organization of the feature space, so that similar patent graphics are concentrated together, further improving the accuracy and efficiency of retrieval.

[0081] In an optional embodiment, the appearance patent graph search network extracts fusion feature vectors, performs clustering division, and constructs a graph search tree, including:

[0082] Based on the existing appearance patent, a multi-view graphic feature extraction network is applied to extract the fused feature vector of the corresponding graphic, and the existing appearance patent, the graphic and the corresponding fused feature vector are stored based on the vector database;

[0083] Based on the two fused feature vectors, the cosine similarity is calculated, the weight is determined, and a similarity matrix is ​​constructed. The fused feature vectors are used as graph nodes and the weights are used as graph edges to construct an undirected weighted graph and start clustering:

[0084] Traversing each graph node in the undirected weighted graph, determining the sum of weights of all graph edges connected to each graph node, constructing a degree matrix, and calculating a Laplace matrix based on the similarity matrix and the degree matrix;

[0085] Performing eigenvalue decomposition on the Laplace matrix to obtain eigenvalues ​​and corresponding eigenvectors, arranging the eigenvalues ​​from small to large, selecting eigenvalues ​​according to a preset feature selection threshold, determining corresponding eigenvectors, constructing a feature matrix, normalizing the feature matrix by rows, and performing clustering based on the K-Means algorithm to obtain a clustering division result;

[0086] The undirected weighted graph is used as the root node of the tree, and the clustering result is the first-layer tree node. Based on each of the first-layer tree nodes, clustering is iteratively performed to obtain the corresponding clustering result, which is used as the subtree node of the first-layer tree node and the second-layer tree node. The subtree nodes are constructed in a loop in sequence until the preset depth upper limit is met to obtain a multi-level index tree. The centrality of each tree node in the index tree is cyclically calculated to determine the center of the tree node.

[0087] The multi-level index tree, combined with the tree node centers, determines a graph search tree.

[0088] The similarity matrix specifically refers to a square matrix in which each element represents the similarity between a pair of objects. Specifically, the degree of similarity between each pair of feature vectors is determined based on the cosine similarity calculated by fusion of feature vectors. In the cluster analysis of appearance patent graphics, the similarity matrix is ​​used to represent the similarity relationship between different patent graphics and is the basis for constructing an undirected weighted graph.

[0089] The degree matrix specifically refers to a diagonal matrix, and the diagonal elements represent the degree of each node in the graph, that is, the sum of the weights of the edges connected to each node. The degree matrix plays an important role in calculating the Laplace matrix and is used to measure the connection strength of each node.

[0090] The Laplace matrix specifically refers to a matrix representation of a graph, which is used for spectral clustering analysis of the graph. It reflects the structural properties of the graph and is constructed by the difference between the degree matrix and the similarity matrix. The Laplace matrix is ​​used for spectral clustering analysis of the graph, and the clustering structure of the graph is identified by eigenvalue decomposition;

[0091] The characteristic matrix specifically refers to a matrix composed of the eigenvectors of the Laplace matrix, where each column is an eigenvector, and the eigenvectors corresponding to the smallest eigenvalues ​​are usually selected. The characteristic matrix is ​​used for K-Means clustering. By normalizing the characteristic matrix by row and performing cluster analysis, the clustering division results of the graph nodes are obtained.

[0092] First, for the existing appearance patent data, the corresponding graphics are processed using a multi-view graphic feature extraction network. By inputting the appearance patent graphics into the pre-trained feature extraction network, a fused feature vector representing the graphics is obtained. The fused feature vector contains the local and global feature information of the graphics at different perspectives and has a strong representation ability. In order to facilitate subsequent retrieval and comparison, the existing appearance patents, the corresponding graphics, and the extracted fused feature vectors are stored in a vector database to form a structured data set.

[0093] Next, in order to measure the similarity between different appearance patent graphics, the similarity between the corresponding fused feature vectors of the appearance patent graphics is calculated. Commonly used similarity measurement methods include Euclidean distance, cosine similarity, etc. Preferably, cosine similarity is used as a metric. By calculating the cosine similarity between each pair of fused feature vectors, a similarity matrix can be obtained, and each element in the matrix represents the degree of similarity between two appearance patent graphics.

[0094] By considering each fused feature vector as a node in the graph and the elements in the similarity matrix as the edge weights between the nodes, a complete graph can be constructed. The connections between the nodes represent the similarity relationship between their corresponding appearance patent graphics. The larger the edge weight, the higher the similarity.

[0095] After obtaining the undirected weighted graph, clustering begins. The first step is to calculate the degree of each node, which is the sum of the weights of all edges connected to the node. By traversing each node in the undirected weighted graph and accumulating the weights of all its connected edges, a degree matrix can be obtained. Each element of the degree matrix represents the degree value of the corresponding node.

[0096] Based on the similarity matrix and degree matrix, the Laplace matrix can be calculated. The Laplace matrix is ​​an important representation of the graph, which contains the topological structure of the graph and the similarity relationship information between nodes. By operating the similarity matrix and degree matrix, the Laplace matrix can be obtained, which provides the basis for subsequent spectral clustering.

[0097] Perform eigenvalue decomposition on the Laplace matrix to obtain a set of eigenvalues ​​and corresponding eigenvectors. Arrange the eigenvalues ​​in ascending order, select the first k smallest eigenvalues ​​according to the preset feature selection threshold, and extract their corresponding eigenvectors. These eigenvectors contain important structural information of the graph and can be used for clustering.

[0098] The selected feature vectors are organized into feature matrices by columns, and L2 normalized by rows so that the L2 norm of each row is 1. The normalized feature matrix can be regarded as the representation of the original fused feature vector in a low-dimensional space, which has better clustering properties.

[0099] Based on the normalized feature matrix, the K-Means algorithm is used for clustering. The K-Means algorithm assigns data points to the nearest cluster center through iterative optimization and continuously updates the cluster center until convergence. By setting the number of clusters K, the appearance patent graphics can be divided into K different clusters, and the graphics within each cluster have a high similarity.

[0100] After obtaining the clustering results of the first layer, the undirected weighted graph is used as the root node of the tree, and each cluster is used as the tree node of the first layer. For each first-layer tree node, the above clustering steps are recursively applied to obtain a finer-grained clustering result as the child node of the node. Similarly, clustering is performed iteratively to construct a multi-level index tree structure. The nodes of each layer represent clusters of different granularities, and the parent node contains the appearance patent graphics represented by the child node.

[0101] In the process of building the index tree, it is necessary to set a depth limit to control the maximum number of layers of the tree. When the preset depth limit is reached, the iteration stops and the final multi-level index tree is obtained.

[0102] In order to further improve the retrieval efficiency, the centrality of each tree node can be calculated. Centrality measures the importance of a node in the graph and can reflect the closeness of the connection between the node and other nodes. Common centrality measures include degree centrality, closeness centrality, etc. By cyclically calculating the centrality of each tree node, the centrality value of each node can be obtained as a reference for retrieval.

[0103] Finally, the constructed multi-level index tree and the centrality information of the tree nodes are combined to form a complete appearance patent graph search tree. The search tree has a hierarchical structure, from the root node to the leaf node, representing clustering divisions of different granularities layer by layer. Through the centrality of the tree nodes, the node most similar to the query graph can be quickly located to improve the retrieval efficiency.

[0104] In this embodiment, the existing appearance patents, graphics and corresponding fused feature vectors are stored in a vector database to achieve efficient management and rapid retrieval of a large amount of data; by calculating the cosine similarity of the fused feature vectors pairwise and constructing a similarity matrix, the similarity between different graphics can be accurately measured, providing a reliable basis for subsequent clustering analysis; the weights are determined according to the cosine similarity, which can better reflect the similarity between different graphics and improve the accuracy of clustering analysis; with the fused feature vectors as nodes, an undirected weighted graph is constructed according to the weights in the similarity matrix to effectively express the association relationship between graphics; a multi-level index tree is constructed based on the clustering division results, which can efficiently manage and organize a large amount of graphic data and improve the efficiency and speed of graphic search; the centrality of each tree node in the index tree is cyclically calculated to determine the center of the tree node, which provides an important reference for constructing a graphic search tree and makes the search results more accurate and targeted.

[0105] S103. Based on the learned graphic similarity measurement network, calculate the similarity score between the feature vector to be applied for and each of the feature vectors of the existing patent, determine the reference graphic set according to the preset selection number, and generate a design patent application recommendation report through a natural language generation model based on the feature vector to be applied for and the feature vectors of the existing patents corresponding to the reference graphic set.

[0106] In this embodiment, the learned graphic similarity measurement network can accurately calculate the similarity score between the feature vector of the patent to be applied for and the feature vector of the existing patent, thereby improving the accuracy and reliability of the similarity measurement; by learning with a large amount of training data, it can adaptively capture the complex relationship between graphic features, thereby improving the generalization ability and precision of the similarity measurement; according to the calculated similarity score, the reference graphic set is determined according to the preset selection number, and the existing patent graphics most similar to the patent to be applied for are effectively screened out to ensure the relevance and representativeness of the reference graphics; by selecting the reference graphic set with the highest similarity, the interference of irrelevant or low-relevant graphics on subsequent processing is reduced, thereby ensuring the quality of the reference data; based on the feature vector to be applied for and the existing patent feature vector corresponding to the reference graphic set, the appearance patent application suggestion report is automatically generated through the natural language generation model, thereby greatly improving the efficiency and accuracy of the report generation; the natural language generation model can generate an application suggestion report with high-quality, structured and detailed content according to the input feature vector, thereby reducing the workload and possible errors of manual writing; the generated application suggestion report has high consistency and standardization, meets the standards and requirements of patent applications, and improves the professionalism and standardization of the application materials.

[0107] In an optional embodiment, based on the learned graph similarity measurement network, the similarity score between the feature vector to be applied for and each feature vector of the existing patent is calculated, and according to a preset selection number, the reference graph set is determined to include:

[0108] Based on the existing appearance patents, a graphic training data set is determined, and a triplet including an anchor sample, a positive sample, and a negative sample is randomly sampled from the graphic training data set, wherein the positive sample and the anchor sample belong to the same tree node in the appearance patent graphic search tree, and the negative sample and the anchor sample do not belong to the same tree node in the appearance patent graphic search tree;

[0109] Based on the multi-view graphic feature extraction network, extracting anchor point feature vectors corresponding to anchor point samples, positive feature vectors corresponding to positive samples, and negative feature vectors corresponding to negative samples respectively;

[0110] Based on the anchor feature vector and the positive feature vector, determine the positive distance between the anchor sample and the positive sample, calculate the positive sample pair loss, based on the anchor feature vector and the negative feature vector, determine the negative distance between the anchor sample and the negative sample, calculate the negative sample pair loss; based on the positive distance and the negative distance, determine the triplet loss function, based on the positive sample pair loss and the negative sample pair loss, determine the contrast loss function; based on the triplet loss function and the contrast loss function, determine the comprehensive loss function;

[0111] Based on the comprehensive loss function, the gradient descent algorithm is used to iteratively optimize the network parameters of the graph similarity measurement network. The network parameters are updated through back propagation to minimize the comprehensive loss until a preset number of iterations is reached to determine the optimal graph similarity measurement network.

[0112] The anchor sample specifically refers to the reference point selected in the triple learning method. In a given data set, the anchor sample is used as a comparison benchmark to compare with other samples. The anchor sample is usually randomly selected in order to learn the characteristics of the data from multiple perspectives.

[0113] The positive sample specifically refers to a sample that is similar to the anchor sample in a specific sense or has the same attributes. In the present invention, the positive sample and the anchor sample belong to the same tree node in the design patent graph search tree, indicating that they are similar in appearance or design, and therefore should be close to each other in the feature space.

[0114] The negative sample specifically refers to a sample that is not similar to the anchor sample in a specific sense or has different attributes. In the present invention, the negative sample and the anchor sample are located at different tree nodes in the design patent graph search tree, indicating that they have significant differences in appearance or design, and therefore should be far apart in the feature space.

[0115] Based on the existing appearance patent data, a graphic training data set is constructed. Random sampling is performed from the data set to form a triplet containing anchor samples, positive samples, and negative samples. Among them, the anchor samples are used as reference samples. The positive samples and anchor samples belong to the same tree node in the appearance patent graphic search tree, indicating that they have high similarity; the negative samples and anchor samples do not belong to the same tree node in the search tree, indicating that they have low similarity. Through this sampling method, sample pairs with different similarity relationships can be obtained, providing effective supervision information for subsequent network training.

[0116] For each triplet, the pre-trained multi-view graphic feature extraction network is used to extract the feature vectors corresponding to the anchor sample, positive sample, and negative sample. By inputting the sample into the feature extraction network, a high-dimensional vector representing the sample feature can be obtained, called the anchor feature vector, positive feature vector, and negative feature vector, which contains the local and global feature information of the sample at different perspectives, providing a basis for similarity measurement.

[0117] Based on the anchor feature vector and the positive feature vector, the distance between the anchor sample and the positive sample is calculated, which is called the positive distance. Common distance measurement methods include Euclidean distance, cosine distance, etc. The smaller the positive distance, the higher the similarity between the anchor sample and the positive sample. At the same time, the positive sample pair loss is calculated based on the positive distance to measure the network's prediction deviation of the similarity of the positive sample pair.

[0118] Similarly, based on the anchor feature vector and the negative feature vector, the distance between the anchor sample and the negative sample is calculated, which is called the negative distance. The larger the negative distance, the lower the similarity between the anchor sample and the negative sample. The negative sample pair loss is calculated based on the negative distance to measure the network's prediction bias for the similarity of the negative sample pair.

[0119] Taking both positive and negative distances into consideration, a triplet loss function is constructed. The triplet loss function aims to minimize the positive distance while maximizing the negative distance, so that the network can learn an effective similarity metric.

[0120] Based on the positive sample pair loss and the negative sample pair loss, a contrast loss function is constructed. The contrast loss function aims to minimize the positive sample pair loss and maximize the negative sample pair loss, so that the network can effectively distinguish sample pairs with different similarities.

[0121] The triplet loss function and the contrast loss function are weighted together to obtain the comprehensive loss function. The comprehensive loss function takes into account both triplet constraints and contrastive learning, and can more comprehensively guide the network to learn an effective similarity measurement model. By adjusting the weights of triplet loss and contrastive loss, the contributions of the two loss functions can be balanced to obtain better network performance.

[0122] Based on the comprehensive loss function, the parameters of the graph similarity measurement network are optimized using the gradient descent algorithm. The gradient descent algorithm calculates the gradient of the loss function to the network parameters, updates the parameters along the negative gradient direction, and gradually minimizes the comprehensive loss.

[0123] In each iteration, the comprehensive loss is calculated through forward propagation, and then the gradient is calculated and the network parameters are updated through back propagation. The back propagation algorithm uses the chain rule to pass the gradient of the loss function to each layer of the network, and adjusts the weights and biases according to the gradient information, so that the output of the network is closer to the desired result.

[0124] Repeat the iterative process and continuously update the network parameters until the preset number of iterations is reached, or the convergence condition is met in advance, and the iteration stops early. During the training process, the generalization ability and optimization progress of the network can be evaluated by monitoring the performance indicators on the validation set, such as accuracy and recall. According to the verification results, hyperparameters such as learning rate and batch size can be appropriately adjusted to further improve network performance.

[0125] After sufficient training and optimization, the optimal graph similarity measurement network is obtained. This network can effectively measure the similarity of appearance patent graphs and correctly distinguish high-similarity positive samples from low-similarity negative samples under given anchor point samples.

[0126] In this embodiment, the distinction between positive samples and negative samples in the triplet provides effective supervision information for network training, which helps the network learn to distinguish between graphics with high similarity and low similarity; the anchor feature vector, positive feature vector and negative feature vector contain rich graphic features, which enhances the accuracy of similarity measurement; by calculating the distance between the anchor sample and the positive sample and the negative sample, the positive distance and negative distance are obtained respectively, which can accurately measure the similarity and difference between samples; the loss is calculated based on the positive distance and the negative distance, and the prediction deviation of the network for the similarity of the positive sample pair and the negative sample pair is measured, which helps to guide network optimization; the triplet loss function effectively constrains the learning process of the network by minimizing the positive distance and maximizing the negative distance, so that the network can The network learns the criteria for distinguishing similar samples from dissimilar samples. The contrast loss function further enhances the network's ability to distinguish sample pairs of different similarities, improving the network's robustness and discrimination. The triplet loss function and the contrast loss function are weightedly combined, and the constraints of the two loss functions are comprehensively considered, so that the network can learn a more comprehensive similarity measurement model. By adjusting the weight of the loss function, the contribution of the triplet loss and the contrast loss can be balanced, further optimizing the network performance. The network parameters are optimized using the gradient descent algorithm, and the parameters are updated along the negative gradient direction by calculating the gradient of the loss function to effectively minimize the comprehensive loss. Through the back-propagation algorithm, the network weights and biases are adjusted layer by layer, so that the network output gradually approaches the expected result.

[0127] In an optional embodiment, the comprehensive loss function has the following formula:

[0128] ;

[0129] in, a represents the anchor point sample, p represents a positive sample, n represents negative samples, L tr represents the triplet loss function, f ( a ) represents the anchor feature vector, f ( p ) represents a positive eigenvector, f ( n ) represents a negative eigenvector, d ( f ( a ), f ( p )) indicates positive distance, d ( f ( a ), f ( n )) represents a negative distance, magin represents the preset minimum distance hyperparameter, L corepresents the contrast loss function, L cb represents the comprehensive loss function, α represents the weight coefficient of the triplet loss function, β Represents the weight coefficient of the contrast loss function.

[0130] The goal of the triplet loss function is to ensure that the distance between the anchor sample and the positive sample is smaller than the distance between the anchor sample and the negative sample, and the difference between the two is at least equal to a preset minimum distance. Specifically, the triplet loss function calculates the distance between the anchor sample and the positive sample, subtracts the distance between the anchor sample and the negative sample, and adds a preset minimum distance. If this value is less than zero, the loss is zero; if greater than zero, the value is used as a loss. By minimizing the triplet loss function, the network is forced to pull the positive sample closer to the anchor sample and push the negative sample away from the anchor sample;

[0131] The goal of the contrastive loss function is to further optimize the similarity metric, which is achieved through two steps. First, the contrastive loss function directly calculates the distance between the anchor sample and the positive sample, and strives to minimize the distance. Secondly, the contrastive loss function imposes an inverse constraint on the distance between the anchor sample and the negative sample: the distance between the anchor sample and the negative sample is calculated, and a preset minimum distance is subtracted. If the result is less than zero, the loss of the corresponding part is zero, otherwise it is the corresponding result, so that the network not only ensures that the positive sample is close to the anchor sample, but also ensures that the negative sample is far away from the anchor sample.

[0132] The comprehensive loss function is a weighted combination of the triplet loss function and the contrast loss function. By giving two weight coefficients, the contribution of the triplet loss and contrast loss to the total loss is adjusted respectively, allowing the network to ensure the proper distribution of positive and negative samples during training, and to improve the network's ability to distinguish between sample similarities and dissimilarities through different constraints; the optimization of the comprehensive loss function can more comprehensively guide the network to learn an effective similarity measurement model.

[0133] In this embodiment, by maximizing the distance difference between the anchor sample and the negative sample, ensuring that the positive sample is close to the anchor sample and the negative sample is far away from the anchor sample, the network can more accurately learn the similarities and dissimilarity between samples, thereby improving the accuracy of the similarity measurement; by simultaneously minimizing the distance between the anchor sample and the positive sample and maximizing the distance between the anchor sample and the negative sample, the network's ability to distinguish sample similarity is further optimized. The introduction of the contrast loss function enables the network to not only focus on the distance between similar sample pairs, but also effectively distinguish sample pairs with different similarities; the comprehensive loss function combines the advantages of triplet loss and contrast loss, provides richer gradient information, can more effectively guide the update of network parameters, and accelerate the convergence speed of training. By considering the constraints of the two loss functions at the same time in each iteration, the network can converge to the optimal solution more quickly.

[0134] In an optional embodiment, randomly sampling from the graphic training data set to construct a triplet including an anchor sample, a positive sample, and a negative sample includes:

[0135] In each training batch, for each anchor sample, calculate the distance between it and all other samples in the batch. Based on the preset positive sample distance threshold, select the sample with the largest distance to the anchor sample as the candidate difficult positive sample. Based on the preset negative sample distance threshold, select the sample with the smallest distance to the anchor sample as the candidate difficult negative sample.

[0136] In multiple training batches, the frequency of positive samples selected as candidate difficult positive samples and the frequency of negative samples selected as candidate difficult negative samples are counted in turn, and the samples are sorted according to the frequency, and the samples with the highest frequency are selected as the final difficult positive samples and the final difficult negative samples;

[0137] Based on the training results of multiple training batches, the positive sample distance threshold and the negative sample distance threshold are dynamically adjusted to construct a triplet containing anchor samples, positive samples and negative samples.

[0138] In each training batch, for each anchor sample in the batch, the distance between the anchor sample and all other samples in the batch is calculated. The distance metric can be a common metric such as Euclidean distance and cosine distance. By calculating the distance between the anchor sample and other samples, the similarity between them can be measured.

[0139] Based on the preset positive sample distance threshold, the sample with the largest distance to the anchor sample is selected in the batch as a candidate difficult positive sample. The positive sample distance threshold is used to determine whether the sample belongs to the positive sample range. Within the positive sample range, the sample with the largest distance is selected as a difficult positive sample. By selecting the sample with the largest distance as a candidate, positive samples with low similarity to the anchor sample can be found, which are challenging for network training.

[0140] Similarly, based on the preset negative sample distance threshold, the sample with the smallest distance to the anchor sample in the batch is selected as a candidate difficult negative sample. The negative sample distance threshold is used to determine whether the sample belongs to the negative sample range. Within the negative sample range, the minimum distance is selected as the difficult negative sample. By selecting the sample with the smallest distance as a candidate, negative samples with high similarity to the anchor sample can be found, which are easily misclassified by the network.

[0141] In multiple training batches, each sample is counted and its frequency of being selected as a candidate difficult positive sample and the frequency of being selected as a candidate difficult negative sample are recorded. The frequency indicates the number of times a sample is selected as a candidate difficult sample in different batches. The higher the frequency, the more challenging the sample is for network training.

[0142] The samples are sorted according to the frequency of positive samples, and the samples with the highest frequency are selected as the final difficult positive samples. The final difficult positive samples are repeatedly selected as candidate difficult positive samples in multiple batches, indicating that they have a low similarity with the anchor samples, which puts higher requirements on the network's discrimination ability. By giving priority to these samples, the network's learning and adaptation capabilities for difficult positive samples can be enhanced.

[0143] Similarly, the samples are sorted according to the frequency of negative samples, and the samples with the highest frequency are selected as the final difficult negative samples. The final difficult negative samples are frequently selected as candidate difficult negative samples in multiple batches, indicating that they are highly similar to the anchor samples, which can easily lead to incorrect judgments by the network. By focusing on these samples, the network's ability to distinguish difficult negative samples can be improved.

[0144] During the training process, the positive sample distance threshold and the negative sample distance threshold are adjusted dynamically. According to the performance of the network in multiple training batches, the size of the threshold is appropriately adjusted to adapt to the learning progress of the network. If the network's ability to distinguish positive samples is improved, the positive sample distance threshold can be appropriately increased to select more challenging difficult positive samples; if the network's ability to distinguish negative samples is enhanced, the negative sample distance threshold can be appropriately lowered to select more confusing difficult negative samples. By dynamically adjusting the threshold, more valuable difficult samples can be continuously mined as the network is optimized, promoting continuous learning and improvement of the network.

[0145] In this embodiment, the network focuses on difficult samples during training, which can significantly improve its feature extraction and matching capabilities, make the extracted features more distinguishable and representative in application scenarios, and improve the performance of appearance patent graphic search and matching; by dynamically adjusting the distance thresholds of positive samples and negative samples, each training batch can select the most challenging positive samples and negative samples, effectively improving the training efficiency, and prompting the network to learn the ability to distinguish similarities and dissimilarities between samples more quickly; samples that are frequently misjudged or difficult to distinguish, namely difficult positive samples and difficult negative samples, are given priority for training, so that the network's ability to distinguish "boundary" samples is significantly enhanced, thereby improving the robustness and generalization ability of the overall model.

[0146] Figure 2 FIG. 1 is a schematic diagram of the structure of an application suggestion system combined with a patent image search according to an embodiment of the present invention. Figure 2 As shown, the system comprises:

[0147] The first unit is used to pre-process the multi-view graphics to be applied for the appearance patent, generate a unified multi-view graphics, extract fusion features based on the unified multi-view graphics through a pre-trained multi-view graphics feature extraction network, and determine the fusion feature vector to be applied for;

[0148] The second unit is used to input the fused feature vector to be applied into the appearance patent graphic search network, search for the closest existing patent, perform multi-view graphic feature extraction on the graphics in the existing appearance patent, and generate the existing patent feature vector, wherein the appearance patent graphic search network extracts the fused feature vector, performs clustering division, and constructs a graphic search tree;

[0149] The third unit is used to calculate the similarity score between the feature vector to be applied for and each of the feature vectors of the existing patent based on the learned graphic similarity measurement network, determine the reference graphic set according to a preset selection number, and generate a design patent application recommendation report based on the feature vector to be applied for and the feature vectors of the existing patents corresponding to the reference graphic set through a natural language generation model.

[0150] According to a third aspect of the embodiments of the present invention,

[0151] An electronic device is provided, comprising:

[0152] processor;

[0153] a memory for storing processor-executable instructions;

[0154] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0155] According to a fourth aspect of the embodiments of the present invention,

[0156] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.

[0157] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.

[0158] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. The application suggestion method combined with the appearance patent image search is characterized by: include: Preprocess the multi-view graphics to be applied for the appearance patent to generate a unified multi-view graphics, extract fusion features based on the unified multi-view graphics through a pre-trained multi-view graphics feature extraction network, and determine the fusion feature vector to be applied for; Input the fused feature vector to be applied into the appearance patent graphic search network, search for the closest existing patent, extract multi-view graphic features of the graphics in the existing patent, and generate the existing patent feature vector, wherein the appearance patent graphic search network extracts the fused feature vector, performs clustering division, and constructs a graphic search tree; Based on the learned graphic similarity measurement network, the similarity score between the feature vector of the patent to be applied for and each of the feature vectors of the existing patents is calculated, and a reference graphic set is determined according to a preset number of selections. Based on the feature vector of the patent to be applied for and the feature vectors of the existing patents corresponding to the reference graphic set, a design patent application recommendation report is generated through a natural language generation model; Preprocessing the multi-view graphics to be applied for the appearance patent to generate a unified multi-view graphics, extracting fusion features based on the unified multi-view graphics through a pre-trained multi-view graphics feature extraction network, and determining the fusion feature vector to be applied for includes: Based on the semantic segmentation method, a binary mask corresponding to the multi-view graphics is generated, the multi-view graphics are segmented, and the target object graphics are extracted; through key point matching, the rotation, translation and scaling of the target object graphics are eliminated, and the target object alignment graphics are determined, and the target object alignment graphics are scaled according to a preset graphic size to generate a unified multi-view graphics; A scale pyramid is constructed based on the unified multi-view graph, each scale graph in the scale pyramid is operated, a local feature corresponding to the scale graph is extracted by a convolution layer to obtain a local feature graph, a transformation layer performs spatial transformation on the local feature graph, a pooling layer performs downsampling on the transformed local feature graph, and a fully connected layer flattens the pooled local feature graph and performs feature transformation and dimension reduction to obtain a multi-scale feature vector; The multi-scale feature vector is input into the feature fusion layer, and the scale weight corresponding to each scale feature is learned iteratively, the attention to each scale feature is adaptively adjusted, and a fused feature vector is generated through element-level weighting.

2. The method according to claim 1, characterized in that The appearance patent graph search network extracts fusion feature vectors, performs clustering division, and constructs a graph search tree, including: Based on the existing appearance patent, a multi-view graphic feature extraction network is applied to extract the fused feature vector of the corresponding graphic, and the existing appearance patent, the graphic and the corresponding fused feature vector are stored based on the vector database; Based on the two fused feature vectors, the cosine similarity is calculated, the weight is determined, and a similarity matrix is ​​constructed. The fused feature vectors are used as graph nodes and the weights are used as graph edges to construct an undirected weighted graph and start clustering: Traversing each graph node in the undirected weighted graph, determining the sum of weights of all graph edges connected to each graph node, constructing a degree matrix, and calculating a Laplace matrix based on the similarity matrix and the degree matrix; Performing eigenvalue decomposition on the Laplace matrix to obtain eigenvalues ​​and corresponding eigenvectors, arranging the eigenvalues ​​from small to large, selecting eigenvalues ​​according to a preset feature selection threshold, determining corresponding eigenvectors, constructing a feature matrix, normalizing the feature matrix by rows, and performing clustering based on the K-Means algorithm to obtain a clustering division result; The undirected weighted graph is used as the root node of the tree, and the clustering result is the first-layer tree node. Based on each of the first-layer tree nodes, clustering is iteratively performed to obtain the corresponding clustering result, which is used as the subtree node of the first-layer tree node and the second-layer tree node. The subtree nodes are constructed in a loop in sequence until the preset depth upper limit is met to obtain a multi-level index tree. The centrality of each tree node in the index tree is cyclically calculated to determine the center of the tree node. The multi-level index tree, combined with the tree node centers, determines a graph search tree.

3. The method according to claim 1, characterized in that Based on the learned graph similarity measurement network, the similarity score between the feature vector of the patent to be applied for and each of the feature vectors of the existing patent is calculated, and according to the preset selection number, the reference graph set is determined to include: Based on the existing appearance patents, a graphic training data set is determined, and a triplet including an anchor sample, a positive sample, and a negative sample is randomly sampled from the graphic training data set, wherein the positive sample and the anchor sample belong to the same tree node in the appearance patent graphic search tree, and the negative sample and the anchor sample do not belong to the same tree node in the appearance patent graphic search tree; Based on the multi-view graphic feature extraction network, extracting anchor point feature vectors corresponding to anchor point samples, positive feature vectors corresponding to positive samples, and negative feature vectors corresponding to negative samples respectively; Based on the anchor feature vector and the positive feature vector, determine the positive distance between the anchor sample and the positive sample, calculate the positive sample pair loss, based on the anchor feature vector and the negative feature vector, determine the negative distance between the anchor sample and the negative sample, calculate the negative sample pair loss; based on the positive distance and the negative distance, determine the triplet loss function, based on the positive sample pair loss and the negative sample pair loss, determine the contrast loss function; based on the triplet loss function and the contrast loss function, determine the comprehensive loss function; Based on the comprehensive loss function, the gradient descent algorithm is used to iteratively optimize the network parameters of the graph similarity measurement network. The network parameters are updated through back propagation to minimize the comprehensive loss until a preset number of iterations is reached to determine the optimal graph similarity measurement network.

4. The method according to claim 3, characterized in that The comprehensive loss function has the following formula: ; in, a represents the anchor point sample, p represents a positive sample, n represents negative samples, L tr represents the triplet loss function, f ( a ) represents the anchor feature vector, f ( p ) represents a positive eigenvector, f ( n ) represents a negative eigenvector, d ( f ( a ), f ( p )) indicates positive distance, d ( f ( a ), f ( n )) represents a negative distance, magin represents the preset minimum distance hyperparameter, L co represents the contrast loss function, L cb represents the comprehensive loss function, α represents the weight coefficient of the triplet loss function, β Represents the weight coefficient of the contrast loss function.

5. The method according to claim 3, characterized in that: Randomly sampling from the graphic training data set to construct a triplet containing an anchor sample, a positive sample, and a negative sample includes: In each training batch, for each anchor sample, calculate the distance between it and all other samples in the batch. Based on the preset positive sample distance threshold, select the sample with the largest distance to the anchor sample as the candidate difficult positive sample. Based on the preset negative sample distance threshold, select the sample with the smallest distance to the anchor sample as the candidate difficult negative sample. In multiple training batches, the frequency of positive samples selected as candidate difficult positive samples and the frequency of negative samples selected as candidate difficult negative samples are counted in turn, and the samples are sorted according to the frequency, and the samples with the highest frequency are selected as the final difficult positive samples and the final difficult negative samples; Based on the training results of multiple training batches, the positive sample distance threshold and the negative sample distance threshold are dynamically adjusted to construct a triplet containing anchor samples, positive samples and negative samples.

6. An application suggestion system combined with a patent image search, used to implement the method described in any one of claims 1 to 5, characterized in that: include: The first unit is used to pre-process the multi-view graphics to be applied for the appearance patent, generate a unified multi-view graphics, extract fusion features based on the unified multi-view graphics through a pre-trained multi-view graphics feature extraction network, and determine the fusion feature vector to be applied for; The second unit is used to input the fused feature vector to be applied into the appearance patent graphic search network, search for the closest existing patent, extract multi-view graphic features of the graphics in the existing patent, and generate the existing patent feature vector, wherein the appearance patent graphic search network extracts the fused feature vector, performs clustering division, and constructs a graphic search tree; The third unit is used to calculate the similarity score between the feature vector of the patent to be applied for and each of the feature vectors of the existing patents based on the learned graphic similarity measurement network, determine the reference graphic set according to a preset selection number, and generate a design patent application recommendation report based on the feature vector of the patent to be applied for and the feature vectors of the existing patents corresponding to the reference graphic set through a natural language generation model.

7. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 5.

8. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Product design patent retrieval analysis system and analysis method thereof

    CN106503144A

  • Deep learning-based multi-view appearance patent image retrieval method

    CN106528826A