Method, non-transitory computer-readable storage medium, and apparatus for searching an image database - Patents.com

Neural networks are used to clean patent images by removing annotations and generating contour masks, addressing inefficiencies in patent image searches and improving retrieval accuracy.

JP7737400B2Active Publication Date: 2025-09-10CAMELOT UK BIDCO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2022567074
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-05-28
Filing Date
2021-05-28
Publication Date
2025-09-10
Estimated Expiration
2041-05-28

AI Technical Summary

Technical Problem

Image searches in patent databases are inefficient and inaccurate due to noise from annotations and descriptive features in patent drawings, which confuse existing image matching algorithms.

Method used

A method involving neural networks to remove noise from patent images by applying a first neural network to identify and remove annotations, followed by a masking process to generate a cleaned image, and using computer vision to detect bounding boxes and generate contour masks, enabling accurate image retrieval.

Benefits of technology

Enables efficient and accurate image retrieval in patent databases by removing noise from patent drawings, allowing for precise identification of similar images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007737400000006
    Figure 0007737400000006
  • Figure 0007737400000007
    Figure 0007737400000007
  • Figure 0007737400000008
    Figure 0007737400000008
Patent Text Reader

Abstract

A method for searching an image database includes receiving a rough image of an object, the rough image including object annotations for visual reference; applying a first neural network to the rough image; correlating results of applying the first neural network with images in a reference database of images, the results including an edited image of the object and images in the reference database of images including the reference object; and selecting one or more images in the reference database of images having a correlation value above a threshold correlation value as matching images. The method can include applying masking to the rough image. The masking can include applying computer vision based object recognition to detect callout features associated with a bounding box and generating a contour mask of the object.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Related Field Cross-References This application is a continuation of U.S. Patent Application No. 17 / 333,707, filed May 28, 2021, and claims the benefit of U.S. Provisional Patent Application No. 63 / 032,432, filed May 29, 2020, each of which is incorporated herein by reference in its entirety for all purposes. [Background technology]

[0002] FIELD This disclosure relates to retrieving images with labels and annotations.

[0003] Performing image searches is complicated by reference and / or sample images that contain image noise, such as annotations and labels. Such image noise can confuse existing image matching algorithms, reducing their output because similar features in the annotations and labels are not matched correctly. This can have a particularly strong impact on patent images and patent image searches.

[0004] Accordingly, the present disclosure addresses these complexities to achieve an image retrieval tool that is both broadly applicable and specifically tailored to searching patent images.

[0005] The foregoing "Background" discussion is intended to generally provide a context for the present disclosure. The inventors' inventions, to the extent described in this Background section, as well as aspects of the description that may not otherwise be considered prior art at the time of filing, are not admitted expressly or impliedly as prior art to the present disclosure. Summary of the Invention [Means for solving the problem]

[0006] FIELD This disclosure relates to searching image databases.

[0007] According to an embodiment, the present disclosure further relates to a method for searching an image database, comprising: receiving, by a processing circuit, a rough image of an object, the rough image including an object annotation for visual reference; applying, by the processing circuit, a first neural network to the rough image; correlating, by the processing circuit, results of applying the first neural network with each image of a reference database of images, the results including each image of the reference database of images including an edited image of the object and the reference object; and selecting, by the processing circuit, one or more images of the reference database of images having a correlation value that exceeds a threshold correlation value as matching images.

[0008] According to an embodiment, the present disclosure further relates to a non-transitory computer-readable storage medium storing computer-readable instructions that, when executed by a computer, cause the computer to perform a method for searching an image database, the method including: receiving a rough image of an object, the rough image including an object annotation for visual reference; applying a first neural network to the rough image; correlating results of applying the first neural network with each image in a reference database of images, the results including an edited image of the object and each image in the reference database of images including the reference object; and selecting as matching images one or more images in the reference database of images having a correlation value that exceeds a threshold correlation value.

[0009] According to an embodiment, the present disclosure further relates to an apparatus for performing a method for searching an image database, the apparatus comprising: a processing circuit configured to: receive a rough image of an object, the rough image including object annotations for visual reference; apply a first neural network to the rough image; correlate results of applying the first neural network with each image in a reference database of images, the results including an edited image of the object and each image in the reference database of images including the reference object; and select, as matching images, one or more images in the reference database of images having a correlation value that exceeds a threshold correlation value.

[0010] The foregoing paragraphs have been provided by way of general description and are not intended to limit the scope of the following claims. The described embodiments, together with further advantages, will be best understood by reference to the following detailed description taken in conjunction with the accompanying drawings, in which: The present invention provides, for example, the following items. (Item 1) 1. A method for searching an image database, comprising: receiving, by a processing circuit, a crude image of an object, the crude image including object annotations for visual reference; applying, by the processing circuitry, a first neural network to the poor image; correlating, by the processing circuitry, a result of applying the first neural network with each image in a reference database of images, the result including an edited image of the object with each image in the reference database of images that includes a reference object; selecting, by said processing circuitry, one or more images of a reference database of said images having a correlation value exceeding a threshold correlation value as matching images; The method comprising: (Item 2) applying the first neural network editing, by the processing circuitry, the poor quality image to remove the object annotation, wherein an output of the editing is the edited image of the object. The method according to item 1. (Item 3) applying, by said processing circuitry, a masking process to said poor image; Item 1, the method of claim 1 further comprising: (Item 4) applying the masking process performing object recognition on the poor quality image by the processing circuitry and via a second neural network to recognize text-based descriptive features; applying computer vision by the processing circuitry and based on the performing the object recognition to detect callout features associated with a bounding box containing the recognized text-based descriptive features; generating, by the processing circuitry and based on the bounding box and the detected callout features, a contour mask of the object; Item 3. The method according to item 3, comprising: (Item 5) editing, by the processing circuitry, the poor quality image based on the result of applying the first neural network and the result of applying the masking process, wherein the result of applying the masking process is the generated contour of the object. The method according to item 4. (Item 6) Item 10. The method of item 1, wherein the first neural network is a convolutional neural network configured to perform pixel-level classification. (Item 7) generating the contour mask of the object, 5. The method of claim 4, further comprising applying object segmentation to the poor quality image by the processing circuitry. (Item 8) 1. A non-transitory computer-readable storage medium storing computer-readable instructions that, when executed by a computer, cause the computer to perform a method for searching an image database, the method comprising: receiving a crude image of an object, the crude image including object annotations for visual reference; applying a first neural network to the poor image; correlating the results of applying the first neural network with each image in a reference database of images, the results including an edited image of the object and each image in the reference database of images that includes a reference object; selecting as matching images one or more images of a reference database of said image having a correlation value exceeding a threshold correlation value; The non-transitory computer-readable storage medium. (Item 9) said applying said first neural network comprises: editing the poor quality image to remove the object annotation, wherein an output of the editing is the edited image of the object. Item 9. The non-transitory computer-readable storage medium of item 8. (Item 10) applying a masking process to said poor image; 9. The non-transitory computer-readable storage medium of item 8, further comprising: (Item 11) applying the masking process performing object recognition on the poor quality image via a second neural network to recognize text-based descriptive features; applying computer vision to detect callout features associated with a bounding box containing the recognized text-based descriptive features based on the performing the object recognition; generating a contour mask of the object based on the bounding box and the detected callout features; Item 11. The non-transitory computer-readable storage medium of item 10, comprising: (Item 12) editing the poor quality image based on the result of applying the first neural network and the result of applying the masking process, wherein the result of applying the masking process is the generated contour of the object. Item 12. The non-transitory computer-readable storage medium of item 11. (Item 13) Item 10. The non-transitory computer-readable storage medium of item 1, wherein the first neural network is a convolutional neural network configured to perform pixel-level classification. (Item 14) generating the contour mask of the object, applying object segmentation to the poor quality image; Item 12. The non-transitory computer-readable storage medium of item 11. (Item 15) 1. An apparatus for performing a method for searching an image database, comprising: a processing circuit; the processing circuitry receiving a crude image of an object, the crude image including object annotations for visual reference; applying a first neural network to the poor image; correlating the results of applying the first neural network with each image in a reference database of images, the results including an edited image of the object and each image in the reference database of images that includes a reference object; selecting as matching images one or more images of a reference database of said image having a correlation value exceeding a threshold correlation value; The apparatus is configured to perform the following: (Item 16) the processing circuitry Item 16. The apparatus of item 15, further configured to apply the first neural network by editing the poor quality image to remove the object annotation, an output of the editing being the edited image of the object. (Item 17) the processing circuitry Item 16. The apparatus of item 15, further configured to apply a masking process to the poor image. (Item 18) the processing circuitry performing object recognition on the poor quality image via a second neural network to recognize text-based descriptive features; applying computer vision to detect callout features associated with a bounding box containing the recognized text-based descriptive features based on the performing the object recognition; generating a contour mask of the object based on the bounding box and the detected callout features; Item 18. The apparatus of item 17, further configured to apply the masking process by (Item 19) the processing circuitry and editing the poor quality image based on the result of applying the first neural network and the result of applying the masking process, wherein the result of applying the masking process is the generated contour mask of the object. Item 19. The device according to item 18. (Item 20) the processing circuitry applying object segmentation to the poor quality image; Item 16. The apparatus of item 15, further configured to generate the contour mask of the object by

[0011] A better understanding of the present disclosure and many of the attendant advantages thereof will be readily obtained as the same becomes more fully understood by reference to the following detailed description when considered in conjunction with the accompanying drawings. [Brief explanation of the drawings]

[0012] [Figure 1] FIG. 1 is a diagram of a mechanical drawing according to an exemplary embodiment of the present disclosure. [Figure 2A] 1 is a flow diagram of a method for searching an image database according to an exemplary embodiment of the present disclosure. [Figure 2B] FIG. 2 is a schematic diagram of a first neural network according to an exemplary embodiment of the present disclosure. [Figure 3] FIG. 10 is a diagram of the results of a first neural network, according to an exemplary embodiment of the present disclosure. [Figure 4] FIG. 10 is a diagram of the results of a first neural network, according to an exemplary embodiment of the present disclosure. [Figure 5] 1 is a flow diagram of a method for searching an image database according to an exemplary embodiment of the present disclosure. [Figure 6A] 10 is a flowchart of a sub-process of a method for searching an image database according to an exemplary embodiment of the present disclosure. [Figure 6B] FIG. 10 is a diagram of the results of a second neural network according to an exemplary embodiment of the present disclosure. [Figure 6C] FIG. 10 is a diagram of a contour mask according to an exemplary embodiment of the present disclosure. [Figure 7A] 10 is a flowchart of a sub-process of a method for searching an image database according to an exemplary embodiment of the present disclosure. [Figure 7B] 1 is a diagram of an edited image according to an exemplary embodiment of the present disclosure. [Figure 8] FIG. 1 is a diagram of a reference drawing used during the training phase, according to an exemplary embodiment of the present disclosure. [Figure 9A] 1 is a flow diagram of a method for synthetically annotating a reference drawing, according to an exemplary embodiment of the present disclosure. [Figure 9B] FIG. 1 is a diagram of a synthetically annotated reference drawing used during the training phase, according to an exemplary embodiment of the present disclosure. [Figure 9C] FIG. 10 is a diagram of the results of a first neural network corresponding to synthetically annotated reference drawings used during the training phase, according to an exemplary embodiment of the present disclosure. [Figure 10A]1 is a flow diagram of a training phase of a neural network deployed in a method for searching an image database, according to an exemplary embodiment of the present disclosure. [Figure 10B] 1 is a flow diagram of a training phase of a neural network deployed in a method for searching an image database, according to an exemplary embodiment of the present disclosure. [Figure 11] 4 is a flow diagram of a training phase of a method for searching an image database according to an exemplary embodiment of the present disclosure. [Figure 12] 1 is a generalized flow diagram of an implementation of an artificial neural network. [Figure 13] 1 is a flow diagram of an implementation of a convolutional neural network, according to an exemplary embodiment of the present disclosure. [Figure 14A] This is an example of a feedforward artificial neural network. [Figure 14B] 1 is an example of a convolutional neural network, according to one embodiment of the present disclosure. [Figure 15] 1 is a schematic diagram of a hardware configuration of a device for performing an image search method according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0013] The terms "a" or "an," as used herein, are defined as one or more. The term "multiple," as used herein, is defined as two or more. The term "another," as used herein, is defined as at least a second or more. The terms "including" and / or "having," as used herein, are defined as comprising (i.e., open language). References throughout this specification to "one embodiment," "a particular embodiment," "one embodiment," "one implementation," "one example," or similar terms mean that a particular feature, structure, or feature described in connection with the embodiment is included in at least one embodiment of the present disclosure. Thus, the appearances of such phrases in various places throughout the specification are not necessarily all referring to the same embodiment. Furthermore, particular features, structures, or features may be combined in any suitable manner in one or more embodiments without limitation.

[0014] Content-based image retrieval, also known as image content query and content-based visual information retrieval, is the application of computer vision techniques to the problem of image retrieval, in other words, the problem of searching for digital images within large databases.

[0015] "Content-based" means that the search analyzes the content of the image rather than metadata, such as keywords, tags, or descriptions, associated with the image. The term "content" in this context can refer to color, shape, texture, or any other information that can be extracted from the image itself. In one example, the term "content" can refer to the entire image. Content-based image retrieval is desirable because searches that rely purely on metadata depend on the quality and completeness of the annotations. To this end, having humans manually annotate images by entering keywords or metadata in a large database can be time-consuming and may not capture the desired keywords to describe the image.

[0016] While any image can be the subject of a content-based image retrieval search, the ability to perform accurate and efficient image searches is crucial for regulatory agencies that must compare images to determine whether similar images have previously been registered. The potential applications of such a system for intellectual property are readily apparent. In particular, such a method could be applied to patent examination, serving as an aid to patent examiners who wish to compare patent drawings, i.e., images, to images in a reference database of patent drawings.

[0017] However, applying image search directly to patent drawings can be inefficient and inaccurate due to "noise," including callouts 101, text-based features, and other descriptive features that merely attempt to point out or describe the underlying objects that are the focus of the patent drawings, as shown in Figure 1. As a result, understanding that similar underlying objects may have different arrangements and noise types, image search applied directly to patent drawings can lead to arriving at inaccurate results in a cumbersome manner.

[0018] In addressing these deficiencies, this disclosure describes methods, computer-readable media, and apparatus for removing "noise" from images, thereby enabling the images to be used for image retrieval. In particular, the methods of this disclosure are applicable to patent drawings, where callouts, text-based features, and other descriptive features may be inconsistent and obscure the underlying objects of the image.

[0019] According to one embodiment, this disclosure describes a method for removing "noise" from patent drawings. The method can be deployed for the generation of a reference image database, or in real time when searching for specific "new" images.

[0020] According to one embodiment, the present disclosure describes a method for searching a database of images, including receiving a poor image of an object and applying a first neural network to the poor image to identify and remove "noise." The method further includes correlating a result of applying the first neural network with each image in a reference database of images, the result including an edited image of the object and each image in the reference database of images that includes the reference object. The method further includes selecting as matching images one or more images in the reference database of images that have a correlation value that exceeds a threshold correlation value.

[0021] According to one embodiment, optical character recognition methods may be integrated with the methods described herein. For example, if reference numbers associated with callouts in patent drawings can be identified, it may be possible to search and find corresponding language in the patent application specification based on the identified reference numbers. Such methods may be included as a pre-processing step to aid in a wide range of search applications.

[0022] Referring now to the drawings in connection with patents, it will be understood that patent specifications often include drawings to explain the subject matter. These drawings contain "noise" such as callouts, text-based features, and other descriptive features. For example, drawings include annotations such as reference numbers, leader lines, and arrows, as shown in FIG. 1. However, such annotations, as introduced above, can interfere with standard image recognition algorithms utilized for image retrieval. Therefore, as described in this disclosure, annotations may be identified and removed to enable the generation of a database of cleaned images or the generation of searchable images in real time.

[0023] 2A and 2B, a method 200 for image retrieval is described, according to one embodiment. It should be understood that method 200 may be performed in part or in whole by local or remote processing circuitry. For example, a single device may perform the steps of method 200, or a first device may send a request to a second device that performs processing.

[0024] At step 205 of method 200, a bad image of an object may be received. A bad image may be an image that has "noise." In the context of a patent application, a bad image may be a patent drawing with annotations such as callouts, text-based features, and other descriptive features. As an example, the patent drawing of FIG. 1 may be considered a bad image.

[0025] In subprocess 210 of method 200, a first neural network may be applied to the poor quality image. In one embodiment, the first neural network may utilize a pixel-level classification model. The neural network may be a convolutional neural network, and the classification model may be trained through deep learning to detect annotations in the image.

[0026] In one embodiment, the first neural network may include at least one encoder / decoder network. The first neural network may be configured to perform a semantic segmentation task. Furthermore, the first neural network may be configured to perform a scene segmentation task by incorporating rich contextual dependencies based on a self-attention mechanism.

[0027] In one embodiment, the first neural network may include two modules, as shown in FIG. 2B and later described with reference to the training stage shown in FIG. 10A . The first module 223 of the first neural network may receive the poor image 222 as input. The first module 223 of the first neural network may include, for example, a convolutional layer and a deconvolutional layer. After processing the poor image 222, the first module 223 of the first neural network may generate an annotation mask 224. The annotation mask 224 may then be used by the second module 226 of the first neural network to identify and remove annotations from the poor image. To this end, the second module 226 of the first neural network may receive the poor image 222 and the annotation mask 224 generated by the first module 223 of the first neural network. The second module 226 of the first neural network may include, for example, a convolutional layer and a deconvolutional layer. The second module 226 of the first neural network may process the poor image 222 taking into account the annotation mask 224 and output an edited image 227 without annotations. Furthermore, the output edited image 227 may include defect inpainting. In other words, in examples where annotations cross the lines of underlying objects in the image and are "inside" the underlying objects, such annotations may be "painted" white to remove the annotations from the image.

[0028] In view of the above, it can be appreciated that the purpose of the first neural network is to remove annotations from a poor image and output a cleaned, edited image. Therefore, the second module of the first neural network 226 can be used alone to output an edited image without input from the first module of the first neural network 223. However, it can also be appreciated that the second module of the first neural network 226 may not be as accurate and / or precise as desired, and thus providing the annotation mask 224 to the second module of the first neural network 226, e.g., as an initial guess, can improve not only the accuracy and / or precision of the second module of the first neural network 226 but also its speed. Furthermore, in connection with training and FIG. 10A , using the first module of the first neural network 223 in combination with the second module of the first neural network 226 can improve both training speed and training accuracy.

[0029] In one embodiment, the first module of the first neural network and the second module of the second neural network can be integrated with a pixel-level segmentation network. For example, the first neural network can include at least one encoder / decoder network combined with a UNet architecture. In one example, the first neural network can be a single neural network. In another example, the modules of the first neural network can be separate neural networks.

[0030] As mentioned above, the output of the first module 223 of the first neural network may be an annotation task, as shown in Figure 3, where the annotations remain but the underlying objects of the bad image are removed. The output of the second module 226 of the first neural network may be a cleaned version of the bad image, as shown in Figure 4, where the identified annotations are removed from the bad image to render the cleaned image. Additionally, empty spaces may be image inpainted for visual fidelity.

[0031] 2A, the cleaned edited image may be processed as part of an image retrieval process in step 245 of method 200. In particular, when the goal is to identify matching images within a reference database of images, step 245 of method 200 may include determining correlations between the cleaned image and each image in the reference database of images.

[0032] At step 250 of method 200, the determined correlations may be evaluated to select at least one matching image. In one embodiment, the selection may be based on a ranking of the determined correlations, where the top n correlated images in the reference database of images are identified. In one embodiment, the selection may be based on a comparison of the determined correlations to a threshold correlation value, where all images in the reference database of images are sufficiently correlated with the selected cleaned image.

[0033] By performing method 200, a patent examiner may be able to, for example, quickly identify several patent drawings and / or patent publications that have images similar to several previously adulterated images.

[0034] 4, or the output of the first neural network of method 200, it can be seen that certain lines, arrows, and numbers remain and may complicate image retrieval. Accordingly, referring now to FIG. 5, a method 500 for image retrieval will be described in accordance with an exemplary embodiment of the present disclosure.

[0035] In one embodiment, method 500 utilizes the first neural network and creation process described with reference to Figures 2A and 2B. The masking process may include computer vision techniques. The masking process, in one example, may include a second neural network. It should be understood that method 500 may be performed in part or in whole by local or remote processing circuitry. For example, a single device may perform each step of method 500, or a first device may send a request to a second device to perform processing.

[0036] At step 505 of method 500, a bad image of an object may be received. A bad image may be an image that has "noise." In the context of a patent application, a bad image may be a patent drawing with annotations such as callouts, text-based features, and other descriptive features. As an example, the patent drawing of FIG. 1 may be considered a bad image.

[0037] At step 510 of method 500, a first neural network may be applied to the poor quality image. In one embodiment, the first neural network may utilize a pixel-level classification model. The neural network may be a convolutional neural network, and the classification model may be trained through deep learning to detect annotations in the image.

[0038] In one embodiment, the first neural network may be a neural network including two modules, as described with reference to Figure 2B. The output of the first module of the first neural network may be a mask of the bad image, as shown in Figure 3, where the annotations remain but the underlying objects in the bad image are removed. The output of the second module of the neural network may be a cleaned version of the bad image, as shown in Figure 4, where the annotation mask from the first module of the first neural network is used to identify and remove the annotations from the bad image to render a cleaned image.

[0039] Simultaneously, in subprocess 515 of method 500, a masking process including a second neural network and computer vision techniques can be applied to the poor quality image. In one embodiment, the second neural network is a label bounding box detector network configured to generate small bounding boxes around reference numbers and figure labels. The detected bounding boxes can then be used with computer vision techniques to detect arrows and lines. Together, the detected bounding boxes and the detected arrows and lines can be used to generate a contour mask that separates the annotations from the underlying objects of the image.

[0040] In particular, referring now to FIGS. 6A-6C, the second neural network may be a convolutional neural network configured to perform object recognition on the poor-quality image in step 620 of sub-process 515. In one example, the second neural network may be a deep learning-based neural network configured to identify reference numbers and generate bounding boxes 617 surrounding the reference numbers and figure labels, as shown in FIG. 6B with corresponding confidence values. In step 625 of sub-process 515, the deep learning-based neural network may be further configured to perform computer vision techniques, either simultaneously or sequentially with the bounding box process, to identify arrows and lines associated with the bounding boxes. Alternatively, step 625 of sub-process 515 may be performed by image processing techniques. Image processing techniques may include, for example, computer vision techniques configured to identify arrows and lines associated with the bounding boxes, either simultaneously or sequentially with the bounding box process. In either case, performing step 625 of sub-process 515 enables the identification of the main outline contours of the underlying objects in the image. Thus, the output of sub-process 515 may be a contour mask 622 associated with the basic objects of the image, generated in step 630 of sub-process 515, as shown in FIG. 6C, which contour mask 622 separates the basic objects of the image.

[0041] 5, the results of applying the first neural network and the results of applying the masking process (i.e., the second neural network and the computer vision technique), either separately or in any combination, can be combined in sub-process 535 of method 500 to edit out the poor image. For example, the edited, cleaned image produced as a result of applying the first neural network can be combined or fused with the contour mask produced in sub-process 515 of method 500.

[0042] To this end, referring now to Figures 7A and 7B, fusion of the results of applying the first neural network and the results of applying the masking process can be performed according to sub-process 535. In step 736 of sub-process 535, the results of applying the first neural network and the output of applying the masking process can be received. The output of the masking process can include detected bounding boxes and a contour mask generated as a result of applying the masking process. At the same time, a poor image can also be received. In step 737 of sub-process 535, the detected bounding boxes are identified in the results of applying the first neural network (i.e., the edited, cleaned image), and the box contents are removed. In step 738 of sub-process 535, the poor image is evaluated to identify groups of connected black pixels as regions. Each region can then be evaluated to determine whether it is outside the contour mask generated by applying the masking process, whether the distance between the region and the nearest bounding box is less than a box distance threshold, and whether the area of ​​the region is less than a region size threshold. If a region in the poor image is determined to meet these three conditions, the region can be eliminated from the edited, cleaned image output as a result of applying the first neural network.

[0043] The edited image generated in sub-process 535 of method 500, an example of which is shown in Figure 7B, can be provided to step 540 of method 500 and used as part of an image retrieval process. In particular, when the goal is to identify matching images within a reference database of images, step 540 of method 200 can include determining a correlation between the edited image and each image in the reference database of images.

[0044] At step 545 of method 500, the determined correlations may be evaluated to select at least one matching image. In one embodiment, the selection may be based on a ranking of the determined correlations, where the top n correlated images in the reference database of images are identified. In one embodiment, the selection may be based on a comparison of the determined correlations to a threshold correlation value, where all images in the reference database of images are sufficiently correlated with the selected edited image.

[0045] In one embodiment, features extracted from either the first neural network, the second neural network, and / or the computer vision technique, separately or in any combination, may also be used separately for image retrieval without the need to generate a combined image. To this end, extracted features of the cleaned image, i.e., the intermediate image, may be used as a searchable image (or feature) for comparison with images in a reference database. In the context of patents, this may enable accurate searching of patent databases to identify patent documents with relevant images. For example, during prosecution of a patent application, an image from a pending application may be used as a query in a search of a patent document reference database, and the results of the search are patent documents containing images that match the query image.

[0046] In other words, according to one embodiment, this disclosure describes a method for searching a patent image database based on a received image input to identify one or more matching patent images. The method may include receiving an image of an object, where the image is, in one example, an annotated image from a patent application; applying a first neural network to the annotated image; generating a clean image of the annotated image based on the application of the first neural network; correlating the clean image of the object with, in one example, images in a reference database of images containing the object or images in a published patent document; and selecting one or more matching images containing the matching object based on the correlation. In one embodiment, the method may further include applying a masking process to the annotated image, where the masking process includes a second neural network configured to detect bounding boxes and computer vision techniques configured to identify arrows, lines, etc. The detected bounding boxes and identified arrows, lines, etc. may be used to generate contour masks of the underlying objects of the image, which are used in combination with the cleaned image generated by applying the first neural network to produce edited cleaned images of the underlying objects of the image.

[0047] In another embodiment, the method may further include extracting features of the object based on the annotation mask, the clean image, and the contour mask, correlating the extracted features of the object with features of each image in a reference database of images containing the object, and selecting one or more matching images containing the matching object based on the correlation.

[0048] Further to the above, it can be appreciated that deep learning typically requires thousands of labeled images to train a classifier (e.g., a convolutional neural network (CNN) classifier). Often, these labeled images are manually labeled, a task that requires significant amounts of time and computational resources. To this end, this disclosure describes a synthetic data generation method that can be configured to process unannotated patent design drawings, such as the drawing in FIG. 8, and apply randomized annotations to the patent design drawings.

[0049] Randomized annotations may be applied so that certain conditions are met, such as ensuring that leader lines do not cross each other and that reference numbers do not overlap on the drawings.

[0050] To this end, Figure 9A provides a method 950 for generating a synthetically annotated drawing. Method 950 is described in the context of generating a single callout, and method 950 can be repeated to annotate images as desired.

[0051] At step 951 of method 950, contour detection is performed to detect lines of basic objects in the image. At step 952 of method 950, a distance map may be generated. The distance map may be a matrix having the same size as the image, with each value in the matrix being the distance to the nearest black pixel. At step 953 of method 950, a random contour pixel may be selected as the starting point of the leader line. In one embodiment, a random perturbation may be added so that the leader line does not always start on a line in the drawing. At step 954 of method 950, an end point of the leader line may be randomly selected. The leader line end points may be biased toward blank areas of the image. In other cases, the selection of leader line end points may be biased toward high values ​​in the matrix (i.e., the distance map). At step 955 of method 950, a check may be performed to identify an intersection between the leader line and an existing leader line. If an intersection is detected, the leader line may be regenerated. Steps 956 and 957 may be performed optionally. At step 956 of method 950, additional control points of the leader line can be sampled to generate a Bezier curve. At step 957 of method 950, an arrowhead can be generated at the start of the leader line. Finally, at step 958 of method 950, random text and fonts can be sampled and placed near the end of the leader line. Additional randomization can be applied, including color shifting and salt and pepper noise.

[0052] Considering method 950, Figure 9B is a diagram of a synthetically annotated image. Figure 8 can be used as a visual reference. Figure 9C provides an isolated view of the synthetic annotations.

[0053] Further to the above, in connection with generating a training database, it can be seen that the present disclosure provides a method for automatically and without human intervention generating a training database for training a neural network.

[0054] Once the training database is synthetically generated, the first and second neural networks described herein can be trained according to the flow diagrams of Figures 10A and 10B. In general, with particular reference to Figure 10A, authentic patent drawings (e.g., mechanical design drawings) can be processed to generate synthetic annotation masks and synthetically annotated images. The synthetically annotated images can be used as training inputs, and the synthetic annotation masks and authentic images can be used as ground truth data. With reference to Figure 10B, synthetically annotated images can be provided as training inputs, and target bounding boxes can be used as ground truth data.

[0055] Referring to FIG. 10A, the first neural network can be divided into a first module 1023 that generates an annotation mask and a second module 1033 that outputs a clean image with annotations removed. Each of the first and second modules can include a convolution block and a deconvolution block. The convolution block, which can include a series of convolution and pooling operations, captures and encodes information in the image data. Gradient backpropagation can be used to learn the weights associated with the convolution / deconvolution process, with the goal of minimizing a loss function (e.g., mean squared error) described below. The annotation mask output from the first module, which can be a binary result, can be evaluated by binary cross-entropy. In this way, it can be seen that the first module is an easier problem to solve (i.e., it converges more quickly). Therefore, its output can be provided to the second module, as shown in FIG. 10A.

[0056] After each model is trained and deployed, patent images are fed in place of synthetically annotated images.

[0057] Referring to FIG. 10B, the second neural network may be an object recognition neural network. In particular, the second neural network may be a convolutional neural network (CNN)-based bounding box detector, which may be trained with synthetic data to identify reference numbers and figure numbers. In one example, a Faster R-CNN architecture may be used, although other architectures may be utilized without departing from the spirit of this disclosure. For example, a first model iteration may be trained with purely synthetic data. This model may then be trained to label real (i.e., non-synthetic) patent drawings, and these labels may be checked by a human reviewer. The human-verified data may then be used to fine-tune the model in a second iteration to improve its accuracy.

[0058] It should be understood that when the first neural network is a pixel-level CNN, the model may be incomplete and may sometimes result in missing annotations. Therefore, this disclosure describes a second neural network as a technique that can be implemented to improve the output of the first neural network. Contour detection, such as that available from the open-source OpenCV library, may be used to find contours in an image, as in FIGS. 6A-6C . Main contours may be detected based on size. Lines that lie outside the main contours and end near the detected reference numbers may be considered leader lines and removed. Similarly, the detected reference numbers may be removed.

[0059] According to one embodiment, a CNN network can be trained on synthetic image data pairs having noisy and noise-free images, and intermediate layers of this trained network can be used as image features to further improve the similarity score, ignoring known noise types.

[0060] Turning now to the figures, Figures 10A and 10B are flow diagrams of a neural network training stage 860 utilized in a method for searching an image database, according to an exemplary embodiment of the present disclosure. In particular, Figures 10A and 10B correspond to the neural networks utilized in method 200 and method 500 of the present disclosure.

[0061] The training stage 860 may include optimization of at least one neural network, which may vary depending on the application and may include a residual network, a convolutional neural network, an encoder / decoder network, etc.

[0062] Generally, each of the at least one network receives as input training data, i.e., for example, synthetically labeled training images. Each of the at least one neural network may be configured to output separate data, but may also output, for example, an estimated image, an estimated annotation mask, or an edited image minimized relative to reference data. In one example, with reference to FIG. 10A , training data 1021 may include a synthetically annotated drawing 1022, a corresponding synthetic annotation mask 1025, and a corresponding cleaned drawing 1027. Training phase 860 may include estimation of an annotation mask 1024 by a first module 1023 of the neural network and estimation of an edited image 1026 by a second module 1033 of the neural network. The estimated drawing 1026 and the estimated annotation mask 1024 may be compared to the cleaned drawing 1027 and the synthetic annotation mask 1025, respectively, to calculate an error between a target value (i.e., the “true” value of the training data) and the estimated value. This error may be evaluated at 1028′ and 1028″, allowing the at least one neural network to iteratively update and improve its predictive power. In another example, with reference to FIG. 10B , the training data 1021 may include synthetically annotated drawings 1022 and corresponding “true” bounding boxes 1037. The “true” bounding boxes 1037 may be human-labeled, e.g., annotations. The training phase 860 may include estimation of the bounding boxes 1029. The estimated bounding boxes 1029 may be compared to the “true” bounding boxes 1037 to calculate the error between the target values ​​(i.e., the “true” values ​​of the training data) and the estimated values. This error may be evaluated at 1028, allowing the at least one neural network to iteratively update and improve its predictive power.

[0063] Specifically, with reference to FIG. 10A, which may be associated with a first neural network of the present disclosure, the training phase 860 may include training a first module 1023 of the first neural network and a second module 1033 of the first neural network.

[0064] Training each module of the first neural network and / or the first neural network as a whole may proceed as described with respect to FIG. 10A.

[0065] In one embodiment, training the first neural network module begins with obtaining training data 1021 from a training database. The training data may include a synthetically annotated drawing 1022 and a corresponding synthetic annotation mask 1025, where the corresponding synthetic annotation mask 1025 is generated based on the synthetically annotated drawing 1022. The synthetically annotated drawing 1022 may be provided as an input layer of the first neural network first module 1023. The input layer may be provided to a hidden layer of the first neural network first module 1023. If the architecture of the neural network first module 1023 follows the convolution / deconvolution description above, the hidden layer may include a contraction stage including one or more of a convolution layer, a concatenation layer, a subsampling layer, a pooling layer, a batch normalization layer, and an activation layer, among others, and a dilation stage including a convolution layer, a concatenation layer, an upsampling layer, an inverse-subsampling layer, and a summation layer, among others. The activation layer may utilize a rectified linear unit (ReLU). The output of the hidden layer becomes the input of the output layer. The output layer, in one example, may be a fully connected layer. The estimated annotation mask 1024 may then be output from the first module 1023 of the first neural network. In one embodiment, the output may be compared to the corresponding synthesized annotation mask 1025 in step 1028′. In practice, the loss function defined therebetween may be evaluated in step 1028′ to determine whether the stopping criteria of the training phase 860 have been met. If the error criteria are met and the loss function is minimized, or if the stopping criteria are otherwise determined to be met, the first module 1023 of the first neural network is deemed sufficiently trained and ready for implementation on unknown real-time data. Alternatively, if the error criteria are not met and the loss function is not minimized, or if the stopping criteria are otherwise determined to be not met, the training phase 860 is repeated, and updates are made to the weights / coefficients of the first module of the first neural network.

[0066] According to one embodiment, since the output from the first module 1023 of the first neural network is binary, the loss function of the first module 1023 of the first neural network is the estimated annotation mask (mask NN ) and the corresponding composite annotation mask (mask synthetic ), i.e., it can be defined as:

number

[0067] Training the first neural network continues with a first neural network second module 1033. To improve training, the output of the first training network first module 1023 can be provided as input to the first neural network second module 1033 along with the synthetically annotated drawing 1022. Training the first neural network second module 1033 begins with obtaining training data from the training database 1021. The training data may include the synthetically annotated drawing 1022 and a corresponding cleaned drawing 1027, where the corresponding cleaned drawing 1027 is a version of the synthetically annotated drawing 1022 without the annotations. The synthetically annotated drawing 1022, along with the estimated annotation mask from the first neural network first module 1023, can be provided as an input layer of the neural network 1023. The input layer can be provided to a hidden layer of the first neural network second module 1033. If the architecture of the first neural network second module 1033 follows the convolution / deconvolution description above, the hidden layer of the first neural network second module 1033 may include a contraction stage including one or more of a convolution layer, a concatenation layer, a subsampling layer, a pooling layer, a batch normalization layer, and an activation layer, among others, and an expansion stage including a convolution layer, a concatenation layer, an upsampling layer, an inverse subsampling layer, and a summation layer, among others. The activation layer may utilize a rectified linear unit (ReLU). The output of the hidden layer of the first neural network second module 1033 becomes the input of the output layer, which may be a fully connected layer, in one example. The edited image 1026 can then be output from the second module 1033 of the first neural network and compared with the corresponding cleaned drawing 1027 in step 1028″. In fact, a loss function defined between them can be evaluated in step 1028″ to determine whether the stopping criterion of the training phase 860 has been met.If it is determined that the error criteria are met and the loss function is minimized, or otherwise the stopping criteria are met, then the second module 1033 of the first neural network is deemed sufficiently trained and ready for implementation on unknown real-time data. Alternatively, if it is determined that the error criteria are not met and the loss function is not minimized, or otherwise the stopping criteria are not met, then the training phase 860 is repeated and updates are made to the weights / coefficients of the second module 1033 of the first neural network.

[0068] According to one embodiment, the loss function of the second module 1033 of the first neural network is a loss function of the edited image (image 1033) output from the second module 1033 of the first neural network. NN ) and the corresponding cleaned image (image cleaned ), that is,

number

[0069] 10B, the training stage 860 of the second neural network 1043 is described. Training the neural network 1023 begins with obtaining training data from the training database 1021. The training data may include synthetically annotated drawings 1022 and corresponding "true" bounding boxes 1037. The "true" bounding boxes 1037 may be human-labeled, e.g., annotations. The synthetically annotated drawings 1022 may be provided as an input layer of the second neural network 1043. The input layer may be provided to a hidden layer of the second neural network 1043. If the architecture of the second neural network 1043 follows the convolution / deconvolution description above, the hidden layer of the second neural network 1043 may include a contraction stage including one or more of a convolution layer, a concatenation layer, a subsampling layer, a pooling layer, a batch normalization layer, and an activation layer, among others, and an expansion stage including a convolution layer, a concatenation layer, an upsampling layer, an inverse subsampling layer, and a summation layer, among others. The activation layer may utilize a rectified linear unit (ReLU). The output of the hidden layer of the second neural network 1043 becomes the input of the output layer, which in one example may be a fully connected layer. A bounding box estimate may be generated at 1029. The estimated bounding box 1029 may be compared to the “true” bounding box 1037 in step 1028 to calculate the error between the target value (i.e., the “true” value of the training data) and the estimated value. This error can be evaluated 1028, allowing the at least one neural network to iteratively update and improve its predictive power. In effect, the loss function defined between them can be evaluated to determine whether the stopping criteria of the training phase 860 have been met. If the error criteria are met and the loss function has been minimized, i.e., otherwise the stopping criteria are determined to be met, the second neural network 1043 is deemed to be sufficiently trained and ready for implementation on unknown real-time data.Alternatively, if it is determined that the error criteria are not met and the loss function is not minimized, i.e., otherwise the stopping criteria are not met, then the training stage 860 is repeated and updates are made to the weights / coefficients of the second neural network 1043.

[0070] According to one embodiment, the loss function of the neural network 1023 is a function of the estimated bounding box image (image 1043) output from the second neural network 1043. NN ) and the corresponding "true" bounding box image (image true ), that is,

number

[0071] The training phase 860 of each of Figures 10A and 10B will now be described with reference to Figures 11 through 14B. It should be understood that although described above as having different loss functions, the neural networks of Figures 10A and 10B may be implemented together during deployment.

[0072] The descriptions of Figures 11 through 14B can be generalized as would be understood by one skilled in the art. Figure 11 shows a flow diagram of an implementation of the training phase performed during optimization of a neural network as described herein.

[0073] During the training phase, representative data from a database of training data is used as training data to train the neural network, resulting in an optimized neural network being output from the training phase. Here, the term "data" may refer to images from a training image database. In an example using training images for data, the training phase of Figures 10A and 10B may be an offline training method that trains the neural network using a large number of synthetic training images that are paired with corresponding cleaned images and annotation masks to train the neural network to estimate the cleaned images and annotation masks, respectively.

[0074] During the training phase, a training database is accessed to obtain multiple databases, and the network is iteratively updated to reduce error (e.g., values ​​generated by a loss function). Updating the network involves iteratively updating, for example, the values ​​of network coefficients at each layer of the neural network, so that the synthetically annotated data processed by the neural network more and more closely matches the target. In other words, the neural network infers the mapping implied by the training data, and a loss function, or cost function, generates an error value related to the discrepancy between the target and the output of the current iteration of the neural network. For example, in a specific implementation, the loss function may use mean squared error to minimize the mean squared error. In the case of a multilayer perceptron (MLP) neural network, a backpropagation algorithm may be used to train the network by minimizing a mean squared error-based loss function using (stochastic) gradient descent. A more detailed description of updating the network coefficients is provided below.

[0075] Training a neural network model essentially means selecting one model from a set of admissible models (i.e., in a Bayesian framework, determining a distribution over the set of admissible models) that minimizes a cost criterion (i.e., an error value calculated using a cost function). In general, neural networks can be trained using any of a number of algorithms for training neural network models (e.g., by applying optimization theory and statistical inference).

[0076] For example, optimization methods used in training neural networks can use a form of gradient descent that incorporates backpropagation to calculate the actual gradient. This is done by taking the derivative of the cost function with respect to the network parameters and then varying those parameters in the gradient-related direction. Backpropagation training algorithms can be steepest descent (e.g., with a variable learning rate, with variable learning rate and momentum, and elastic backpropagation), quasi-Newton methods (e.g., Broyden-Fletcher-Goldfarb-Shanno, one-step secant, and Levenberg-Marquardt), or conjugate gradient methods (e.g., Fletcher-Reeves update, Polak-Ribiere update, Powell-Beale restart, and scaled conjugate gradient). Additionally, evolutionary methods such as gene expression programming, simulated annealing, expectation maximization, nonparametric methods, and particle swarm optimization can also be used to train neural networks.

[0077] 11, the flow diagram is a non-limiting example of an implementation of the training stage 860 for training a neural network using training data. The data for the training data can be from any training data set in the training database.

[0078] In step 1180 of the training phase 860, initial guesses are generated for the neural network's coefficients. For example, the initial guesses may be based on one of LeCun initialization, Xavier initialization, and Kaiming initialization.

[0079] Steps 1181 through 860 provide a non-limiting example of an optimization method for training a neural network. In step 1181 of the training phase 860, an error is calculated (e.g., using a loss function / cost function) to represent a measure of difference (e.g., a distance measure) between the target and the neural network output data applied in the current iteration of the neural network. The error can be calculated using any known cost function or distance measure between image data, including those cost functions described above. Furthermore, in certain implementations, the error / loss function can be calculated using one or more of hinge loss and cross-entropy loss.

[0080] Furthermore, loss functions can be combined with regularization techniques to avoid overfitting the network to the specific examples represented in the training data. Regularization can be useful for preventing overfitting in machine learning problems. If training takes too long and the model is assumed to be sufficiently expressive, the network will learn the noise specific to that dataset, known as overfitting. In the case of overfitting, the neural network will generalize poorly and the noise will vary across datasets, resulting in high variance. The minimum total error occurs when the sum of bias and variance is minimized. Therefore, it is desirable for the trained network to reach a local minimum that describes the data in the simplest way possible to maximize the likelihood that it represents a general solution rather than a solution specific to the noise in the training data. This goal can be achieved, for example, by early stopping, weight regularization, lasso regularization, ridge regularization, or elastic net regularization.

[0081] In a specific implementation, a neural network is trained using backpropagation. Backpropagation can be used to train a neural network, in conjunction with gradient descent optimization. During the forward pass, the algorithm calculates predictions for the network based on the current parameters θ. These predictions are then input into a loss function, which compares the predictions with the corresponding ground truth labels. During the backward pass, the model calculates the gradient of the loss function with respect to the current parameters, and then the parameters are updated by taking a predefined step size in the direction of minimized loss (e.g., in accelerated methods such as Nesterov momentum methods and various adaptive methods, the step size can be selected to converge more quickly to optimize the loss function).

[0082] The optimization method in which the backprojection is performed can use one or more of gradient descent, batch gradient descent, stochastic gradient descent, and mini-batch stochastic gradient descent. Additionally, the optimization method can be accelerated using one or more momentum update techniques of optimization techniques that result in faster convergence rates of stochastic gradient descent in deep networks, including adaptive methods including, for example, the Nesterov momentum technique or the Adagrad subgradient method, the Adadelta or RMSProp parameter update variants of the Adagrad method, and the Adam adaptive optimization technique. The optimization method can also apply a quadratic method by incorporating a Jacobian matrix into the update step.

[0083] Forward and backward passes can be performed incrementally through each layer of the network. In the forward pass, execution begins by feeding inputs through the first layer, thus causing activations in the outputs of subsequent layers. This process is repeated until the loss function of the last layer is reached. During the backward pass, the last layer calculates gradients with respect to its own learnable parameters (if any) and also with respect to its own input, which serves as the upstream derivative of the previous layer. This process is repeated until the input layer is reached.

[0084] Returning to the non-limiting example shown in Figure 11, step 1182 of training phase 860 determines the change in error, such that a function of the change in the network can be calculated (e.g., an error gradient), and this change in error can be used to select the direction and step size of subsequent changes in the neural network's weights / coefficients. Calculating the error gradient in this manner is consistent with certain implementations of gradient descent optimization methods. In certain other implementations, this step can be omitted and / or replaced with another step in accordance with a different optimization algorithm (e.g., a non-gradient descent optimization algorithm such as simulated annealing or a genetic algorithm), as will be understood by those skilled in the art.

[0085] In step 1183 of the training phase 860, a new set of coefficients is determined for the neural network. For example, the weights / coefficients may be updated using the changes calculated in step 1182, as in a gradient descent optimization method or an over-relaxation accelerated method.

[0086] In step 1184 of the training phase 860, new error values ​​are calculated using the updated weights / coefficients of the neural network.

[0087] In step 1185 of the training phase 860, a predefined stopping criterion is used to determine whether training of the network is complete. For example, the predefined stopping criterion may evaluate whether the new error and / or the total number of iterations performed exceeds a predefined value. For example, the stopping criterion may be met if either the new error falls below a predefined threshold or a maximum number of iterations has been reached. When the stopping criterion is not met, the training phase 860 returns and continues to the beginning of the iterative loop by repeating step 1182 using the new weights and coefficients (the iterative loop includes steps 1182, 1183, 1184, and 1185). Once the stopping criterion is met, the training phase 860 is complete.

[0088] 12 and 13 illustrate flow diagrams of neural network implementations, aspects of which may be incorporated into the training and / or runtime stages of the neural networks of the present disclosure. FIG. 12 is general to any type of layer in a feedforward artificial neural network (ANN), including, for example, a fully connected layer. FIG. 13 is specific to the convolutional, pooling, batch normalization, and ReLU layers of a CNN. A P3DNN may include aspects of the flow diagrams of FIGS. 12 and 13, including fully connected, convolutional, pooling, batch normalization, and ReLU layers, as would be understood by one skilled in the art.

[0089] In step 1287 of the training phase 860, weights / coefficients corresponding to connections between neurons (ie, nodes) are applied to each input corresponding to, for example, pixels of a training image.

[0090] The weighted inputs are summed in step 1288. As the only non-zero weights / coefficients connecting a given neuron in the next layer are partially localized in the image represented in the previous layer, the combination of steps 1287 and 1288 is essentially equivalent to performing a convolution operation.

[0091] In step 1289, each threshold is applied to the weighted sum of each neuron.

[0092] In sub-process 1290, the weighting, summing, and thresholding steps are repeated for each subsequent layer.

[0093] Figure 13 shows a flow diagram of another implementation of a neural network. The implementation of the neural network shown in Figure 13 corresponds to operating on training images in a hidden layer using a non-limiting implementation of a neural network.

[0094] In step 1391, the computation for the convolutional layer is performed as described above and in accordance with the understanding of convolutional layers by those skilled in the art.

[0095] In step 1392, following the convolution, batch normalization can be performed to control for variability in the output of the previous layer, as will be understood by those skilled in the art.

[0096] In step 1393, following batch normalization, activation is performed in light of the above description of activation and in accordance with the understanding of activation by those skilled in the art. In one example, the activation function is a modified activation function, or for example, ReLU, as described above.

[0097] In another implementation, the ReLU layer of step 1393 may be performed before the batch normalization layer of step 1392.

[0098] In step 1394, the output from the convolutional layer, followed by batch normalization and activation, is the input to a pooling layer, which is performed in light of the above description of pooling layers and in accordance with the understanding of pooling layers by those skilled in the art.

[0099] In process 1395, the steps of convolutional layers, pooling layers, batch normalization layers, and ReLU layers may be repeated in whole or in part for a predefined number of layers. Following (or intermixed with) the above-mentioned layers, the output from the ReLU layer may be fed to a predefined number of ANN layers performed according to the description provided for the ANN layers of FIG. 11.

[0100] 14A and 14B show various examples of interconnections between layers within a neural network. A neural network can include, by way of example, fully connected layers, convolutional layers, subsampling layers, concatenation layers, pooling layers, batch normalization layers, and activation layers, all of which are described above and below. In a particular preferred implementation of a neural network, convolutional layers are placed near the input layer, while fully connected layers, which perform high-level inference, are placed further down in the architecture toward the loss function. Pooling layers are inserted after convolutions and can provide reduction, reducing the spatial range of the filters and, therefore, the amount of learnable parameters. Batch normalization layers adjust for gradient perturbations to outliers and accelerate the learning process. Activation functions are also incorporated into various layers to introduce nonlinearity and enable the network to learn complex predictive relationships. Activation functions can be saturating activation functions (e.g., sigmoid or hyperbolic tangent activation functions) or modified activation functions (e.g., the ReLU activation functions described above).

[0101] Figure 14A shows an example of a generic ANN with N inputs, K hidden layers, and three outputs. Each layer is composed of nodes (also called neurons), which perform a weighted sum of their inputs, compare the result of the weighted sum to a threshold, and generate an output. ANNs constitute a class of functions whose membership is obtained by varying architectural details such as thresholds, connection weights, or the number of nodes and / or their connectivity. Nodes in an ANN are sometimes called neurons (or neuronal nodes), and neurons can have interconnections between different layers of the ANN system. The simplest ANN has three layers and is called an autoencoder. Neural networks have three or more neuron layers, with as many output neurons as input neurons.

number

[0102] Mathematically, the network function m(x) of neurons can be further defined as a composition of other functions n i (x). This can be conveniently represented as a network structure, with arrows indicating dependencies between variables, as shown in Figures 14A and 14B. For example, an ANN can use a nonlinear weighted sum, where m(x) = K(Σ i w i n i (x)), and K (commonly called the activation function) is some predefined function such as the hyperbolic tangent.

[0103] In Figure 14A (and similarly in Figure 14B), neurons (i.e., nodes) are represented by circles around threshold functions. In the non-limiting example shown in Figure 14A, inputs are represented as circles around linear functions, and arrows indicate directed communication between neurons. In a particular implementation, the neural network is a feedforward network.

[0104] The P3DNN of this disclosure searches within a class of functions F and uses a set of observations to learn and solve a particular task in some optimal sense (e.g., a stopping criterion). * It operates to accomplish a specific task, such as estimating a cleaned image, by finding a solution m * in the case of,

number

[0105] FIG. 14B shows a non-limiting example in which the neural network is a CNN. A CNN is a type of ANN that has properties beneficial to image processing and therefore has particular relevance to image denoising applications. CNNs use feedforward ANNs, where the pattern of connectivity between neurons can represent convolutions in image processing. For example, CNNs can be used for image processing optimization by using multiple layers of small neuron ensembles that process portions of the input image, called receptive fields. The outputs of these ensembles can then be tiled so that they overlap to obtain a better representation of the original image. This processing pattern can be repeated over multiple layers with convolutional and pooling layers, as shown, and can include batch normalization and activation layers.

[0106] As generally applied above, following the convolutional layers, the CNN may include local and / or global pooling layers that combine the outputs of the neuron clusters in the convolutional layers. Furthermore, in particular implementations, the CNN may also include various combinations of convolutional and fully connected layers, with pointwise nonlinearities applied at the end of or after each layer.

[0107] Next, a hardware description of an apparatus, i.e., a device, for retrieving images according to an exemplary embodiment will be described with reference to FIG. 15. In FIG. 15, the device includes a CPU 1500 that executes the processes described above / below. Process data and instructions may be stored in memory 1502. These process and instructions may be stored on a storage medium disk 1504, such as a hard drive (HDD) or a portable storage medium, or may be stored remotely. Furthermore, the claimed advancement is not limited by the form of the computer-readable medium on which the instructions for the inventive processes are stored. For example, the instructions may be stored on a CD, DVD, flash memory, RAM, ROM, PROM, EPROM, EEPROM, hard disk, or any other information processing device with which a device, such as a server or computer, communicates.

[0108] Additionally, the claimed advancements may be provided as utility applications, background daemons, or operating system components, or combinations thereof, that run in conjunction with CPU 1500 and an operating system such as Microsoft Windows® 7, Microsoft Windows® 10, UNIX®, Solaris, LINUX®, Apple MAC-OS, and other systems known to those skilled in the art.

[0109] The hardware elements for achieving the device may be realized by various circuit elements known to those skilled in the art. For example, CPU 1500 may be an Intel Xenon or Core processor, or an AMD Opteron processor, or other processor types that will be recognized by those skilled in the art. Alternatively, CPU 1500 may be implemented on an FPGA, an ASIC, a PLD, or using discrete logic circuits, as will be recognized by those skilled in the art. Furthermore, CPU 1500 may be implemented as multiple processors cooperating in parallel to execute instructions of the inventive process described above.

[0110] The device of Figure 15 also includes a network controller 1506, such as an Intel Ethernet PRO network interface card from Intel Corporation of America, for interfacing with a network 1555. As can be appreciated, the network 1555 can be a public network, such as the Internet, or a private network, such as a LAN or WAN network, or any combination thereof, and can also include PSTN or ISDN subnetworks. The network 1555 can also be wired, such as an Ethernet network, or wireless, such as a cellular network, including EDGE, 3G, and 4G wireless cellular systems. The wireless network can also be WiFi, Bluetooth, or any other known wireless form of communication.

[0111] The device further includes a display controller 1508, such as an NVIDIA GeForce GTX or Quadro graphics adapter from NVIDIA Corporation of America, for interfacing with a display 1510, such as a Hewlett Packard HPL2445w LCD monitor. A general-purpose I / O interface 1512 interfaces with a keyboard and / or mouse 1514, and a touchscreen panel 1516 on or separate from the display 1510. The general-purpose I / O interface also connects to a variety of peripherals 1518, including printers and scanners, such as Hewlett Packard's Officelet or DeskJet.

[0112] A sound controller 1520, such as Creative's Sound Blaster X-Fi Titanium, is also provided with the device to interface with a speaker / microphone 1522 and thereby provide sound and / or music.

[0113] A generic storage controller 1524 connects the storage media disk 1504 to a communication bus 1526, which may be ISA, EISA, VESA, PCI, or the like, for interconnecting all of the device's components. A description of the general features and functionality of the display 1510, keyboard and / or mouse 1514, as well as the display controller 1508, storage controller 1524, network controller 1506, sound controller 1520, and generic I / O interface 1512 is omitted herein for the sake of brevity, as these features are known.

[0114] Naturally, many modifications and variations are possible in light of the above teachings, and it is therefore to be understood that within the scope of the appended claims, embodiments of the present disclosure may be practiced other than as specifically described herein.

[0115] Embodiments of the present disclosure may also be as shown in the following parentheses:

[0116] (1) A method for searching an image database, the method comprising: receiving, by a processing circuit, a rough image of an object, the rough image including object annotations for visual reference; applying, by the processing circuit, a first neural network to the rough image; correlating, by the processing circuit, a result of applying the first neural network with each image in a reference database of images, the result including an edited image of the object and each image in the reference database of images including a reference object; and selecting, by the processing circuit, one or more images in the reference database of images having a correlation value that exceeds a threshold correlation value as matching images.

[0117] (2) The method of (1), wherein applying the first neural network includes editing, by the processing circuit, the poor quality image to remove the object annotation, wherein an output of the editing is the edited image of the object.

[0118] (3) The method of either (1) or (2), further comprising applying a masking process to the poor image by the processing circuitry.

[0119] (4) A method according to any one of (1) to (3), wherein applying the masking process includes performing object recognition on the poor quality image by the processing circuit and via a second neural network to recognize text-based descriptive features; applying computer vision by the processing circuit and based on the performing of the object recognition to detect callout features associated with a bounding box containing the recognized text-based descriptive features; and generating a contour mask of the object by the processing circuit and based on the bounding box and the detected callout features.

[0120] (5) A method according to any one of (1) to (4), further comprising editing the poor quality image by the processing circuit based on the result of applying the first neural network and the result of applying the masking process, wherein the result of applying the masking process is the generated contour of the object.

[0121] (6) The method according to any one of (1) to (5), wherein the first neural network is a convolutional neural network configured to perform pixel-level classification.

[0122] (7) A method according to any one of (1) to (6), wherein generating the contour mask of the object includes applying object segmentation to the poor quality image by the processing circuitry.

[0123] (8) A non-transitory computer-readable storage medium storing computer-readable instructions that, when executed by a computer, cause the computer to perform a method for searching an image database, the method including: receiving a rough image of an object, the rough image including object annotations for visual reference; applying a first neural network to the rough image; correlating a result of applying the first neural network with each image in a reference database of images, the result including an edited image of the object and each image in the reference database of images including the reference object; and selecting, as matching images, one or more images in the reference database of images having a correlation value that exceeds a threshold correlation value.

[0124] (9) The non-transitory computer-readable storage medium of (8), wherein applying the first neural network includes editing the poor quality image to remove the object annotation, and an output of the editing is the edited image of the object.

[0125] (10) The non-transitory computer-readable storage medium of either (8) or (9), further comprising applying a masking process to the poor image.

[0126] (11) A non-transitory computer-readable storage medium according to any one of (8) to (10), wherein applying the masking process includes performing object recognition on the poor quality image via a second neural network to recognize text-based descriptive features; applying computer vision based on the performing object recognition to detect callout features associated with a bounding box containing the recognized text-based descriptive features; and generating a contour mask of the object based on the bounding box and the detected callout features.

[0127] (12) A non-transitory computer-readable storage medium according to any one of (8) to (11), further comprising editing the poor quality image based on the result of applying the first neural network and the result of applying the masking process, wherein the result of applying the masking process is the generated contour of the object.

[0128] (13) The non-transitory computer-readable storage medium of any one of (8) to (12), wherein the first neural network is a convolutional neural network configured to perform pixel-level classification.

[0129] (14) A non-transitory computer-readable storage medium according to any one of (8) to (13), wherein generating the contour mask of the object includes applying object segmentation to the rough image.

[0130] (15) An apparatus for performing a method for searching an image database, the apparatus comprising: a processing circuit configured to: receive a rough image of an object, the rough image including an object annotation for visual reference; apply a first neural network to the rough image; correlate a result of applying the first neural network with each image in a reference database of images, the result including an edited image of the object and each image in the reference database of images including the reference object; and select as matching images one or more images in the reference database of images having a correlation value that exceeds a threshold correlation value.

[0131] (16) The apparatus of (15), wherein the processing circuitry is further configured to apply the first neural network by editing the poor quality image to remove the object annotation, an output of the editing being the edited image of the object.

[0132] (17) The apparatus of either (15) or (16), wherein the processing circuitry is further configured to apply a masking process to the poor image.

[0133] (18) The device described in any one of (15) to (17), wherein the processing circuitry is further configured to perform object recognition on the poor quality image via a second neural network to recognize text-based descriptive features, apply computer vision based on the performing of the object recognition to detect callout features associated with a bounding box containing the recognized text-based descriptive features, and apply the masking process by generating a contour mask of the object based on the bounding box and the detected callout features.

[0134] (19) The device described in any one of (15) to (18), wherein the processing circuit is further configured to edit the poor quality image based on the result of applying the first neural network and the result of applying the masking process, wherein the result of applying the masking process is the generated contour mask of the object.

[0135] (20) The device described in any one of (15) to (19), wherein the processing circuitry is further configured to generate the contour mask of the object by applying object segmentation to the rough image.

[0136] Accordingly, the foregoing description discloses and describes merely exemplary embodiments of the present disclosure. As will be understood by those skilled in the art, the present disclosure may be embodied in other specific forms without departing from its spirit or essential characteristics. Accordingly, the disclosure of the present disclosure is intended to be illustrative and not limiting of the scope of the present disclosure and other claims. The disclosure includes readily identifiable variations of the teachings herein, and in part defines the scope of the following claim terms so as not to subject the subject matter to the public.

Claims

1. 1. A method for searching an image database, comprising: receiving, by a processing circuit, a crude image of an object, the crude image including object annotations for visual reference; applying, by the processing circuitry, a first neural network to the poor image; correlating, by the processing circuitry, results of applying the first neural network with each image in a reference database of images, the results including an edited image of the object and each image in the reference database of images that includes a reference object; selecting, by said processing circuitry, one or more images of a reference database of said images having a correlation value exceeding a threshold correlation value as matching images; applying, by the processing circuitry, a masking process to the poor image, wherein applying the masking process includes: performing object recognition on the poor quality image by the processing circuitry and via a second neural network to recognize text-based descriptive features; applying computer vision by the processing circuitry and based on performing the object recognition to detect callout features associated with a bounding box containing the recognized text-based descriptive features; generating, by the processing circuitry and based on the bounding box and the detected callout features, a contour mask of the object; Including A method comprising:

2. applying the first neural network; editing, by the processing circuitry, the poor quality image to remove the object annotation, an output of said editing being the edited image of the object. The method of claim 1.

3. editing, by the processing circuitry, the poor quality image based on the result of applying the first neural network and the result of applying the masking process, wherein the result of applying the masking process is the generated contour mask of the object. The method of claim 1.

4. The method of claim 1 , wherein the first neural network is a convolutional neural network configured to perform pixel-level classification.

5. generating the contour mask of the object, The method of claim 1 , further comprising applying, by the processing circuitry, object segmentation to the poor quality image.

6. A non-transitory computer-readable storage medium storing computer-readable instructions that, when executed by a computer, cause the computer to perform a method for searching an image database, the method comprising: receiving a crude image of an object, the crude image including object annotations for visual reference; applying a first neural network to the poor image; correlating the results of applying the first neural network with each image in a reference database of images, the results including an edited image of the object and each image in the reference database of images that includes a reference object; selecting as matching images one or more images of a reference database of said image having a correlation value above a threshold correlation value; applying a masking process to the poor image, wherein applying the masking process comprises: performing object recognition on the poor quality image via a second neural network to recognize text-based descriptive features; Based on performing the object recognition, applying computer vision to detect callout features associated with a bounding box containing the recognized text-based descriptive features; generating a contour mask of the object based on the bounding box and the detected callout features; Including 1. A non-transitory computer-readable storage medium comprising:

7. applying the first neural network; editing the poor quality image to remove the object annotation, wherein an output of said editing is the edited image of the object. The non-transitory computer-readable storage medium of claim 6.

8. editing the poor quality image based on the result of applying the first neural network and the result of applying the masking process, wherein the result of applying the masking process is the generated contour mask of the object. The non-transitory computer-readable storage medium of claim 6.

9. 7. The non-transitory computer-readable storage medium of claim 6, wherein the first neural network is a convolutional neural network configured to perform pixel-level classification.

10. generating the contour mask of the object, applying object segmentation to the poor quality image; The non-transitory computer-readable storage medium of claim 6.

11. 1. An apparatus for performing a method for searching an image database, the apparatus comprising: a processing circuit; the processing circuitry receiving a crude image of an object, the crude image including object annotations for visual reference; applying a first neural network to the poor image; correlating the results of applying the first neural network with each image in a reference database of images, the results including an edited image of the object and each image in the reference database of images that includes a reference object; selecting as matching images one or more images of a reference database of said image having a correlation value above a threshold correlation value; applying a masking process to the poor image; and the processing circuitry is configured to: performing object recognition on the poor quality image via a second neural network to recognize text-based descriptive features; Based on performing the object recognition, applying computer vision to detect callout features associated with a bounding box containing the recognized text-based descriptive features; generating a contour mask of the object based on the bounding box and the detected callout features; an apparatus configured to apply the masking process by

12. the processing circuitry 12. The apparatus of claim 11, further configured to apply the first neural network by editing the poor image to remove the object annotation, an output of said editing being the edited image of the object.

13. the processing circuitry and further configured to edit the poor quality image based on the result of applying the first neural network and the result of applying the masking process, wherein the result of applying the masking process is the generated contour mask of the object.

12. The apparatus of claim 11.

14. the processing circuitry applying object segmentation to the poor quality image; The apparatus of claim 11 , further configured to generate the contour mask of the object by:

Citation Information

Patent Citations

  • Image-based document indexing and retrieval

    JP2005251169A

  • Apparatus component image search device for assembly drawing

    JP2006113922A

  • Method and system for extracting information from hand-marked industrial inspection sheet

    JP2019087222A

  • Image-based document indexing and retrieval

    US20050165747A1

  • Device part assembly drawing image search apparatus

    US20060082595A1