System and method for part identification and evaluation using multiple images

By training a neural network to observe parts from multiple angles and conditions, combined with similarity calculation and probability combination, the problems of inaccurate and high cost of part identification in existing technologies are solved, and more efficient part identification and ordering recommendations are achieved.

CN116615766BActive Publication Date: 2025-09-23CATERPILLAR INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202180084097.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-12-15
Filing Date
2021-12-08
Publication Date
2025-09-23
Estimated Expiration
2041-12-08

AI Technical Summary

Technical Problem

Existing technologies have difficulty in accurately identifying worn parts and suffer from problems of part label confusion and increased costs. Especially in the image recognition process, existing methods increase the complexity of the image training set.

Method used

By using multiple images for object recognition, training the neural network to observe parts from different angles and conditions, and combining similarity calculation and probabilistic combination, recognition accuracy is improved.

Benefits of technology

Improves part identification accuracy and efficiency, reduces misidentifications, lowers costs, and provides part information and ordering recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116615766B_ABST
    Figure CN116615766B_ABST
Patent Text Reader

Abstract

A method (100) for object recognition using multiple images. The method includes training an object recognition model (102). Training the model includes collecting training images (202) of each of a plurality of objects, labeling each of the plurality of training images with a corresponding one of a plurality of object identifiers (206), and training a neural network (208) using the plurality of labeled training images. At least two target images of a target object are received and fed into the trained object recognition model (104). The method also includes receiving, from the trained object recognition model, an object identifier corresponding to the target object and a probability that the object identifier corresponds to the target object for each of the at least two target images (106). A similarity value (108) is calculated between the at least two target images, and the probabilities (110) of the at least two target images are combined in proportion to the similarity value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This patent application relates to machine maintenance, and more particularly to part identification and evaluation. Background Art

[0002] As equipment is used, certain parts gradually wear out and should be replaced when a part fails or becomes worn to the point where it begins to degrade the equipment's performance. Sometimes it's difficult to identify worn parts because certain identifying features may have worn off, the part number may be obscured by other parts on the machine, and / or the part may be dirty. Parts can also be mislabeled during the manufacturing process, leading to confusion and increased costs.

[0003] For example, efforts have been made to use image recognition to identify parts to verify that the correct part is used for replacement and to verify that parts received from suppliers are correctly labeled. Image recognition can also be used to detect defects in new and used parts. Image recognition technology is known, and some methods have been developed to identify objects, such as parts.

[0004] For example, U.S. Patent No. 10,664,722 to Sharma et al. (hereinafter referred to as "Sharma") describes a method for training a neural network to classify a large number of objects that may be encountered, for example, in a supermarket. The method includes assigning training images to a plurality of buckets or pools corresponding thereto, the assignment including assigning training images that exemplify a first object class to a first bucket and assigning training images that exemplify a second object class to a second bucket. The method also includes finding a hotspot excerpt within a training image depicting an object of a particular object class that is more important to the neural network's classification of the image than another excerpt within the training image. A new image is created by overlaying a copy of the hotspot excerpt on a background image and assigning the new image to a bucket that is not associated with the particular object class. The new image is presented to the neural network during a training cycle, wherein presenting the new image to the neural network has the effect of reducing the importance of the hotspot excerpt in classifying images depicting objects of the particular object class.

[0005] As another example, U.S. Patent Publication No. 2019 / 0392318 to Ghafoor et al. (referred to as "Ghafoor") describes a system for training a computational neural network to recognize objects and / or actions from images. The system includes a training unit configured to receive a plurality of images captured from one or more cameras, each image having an associated timestamp indicating the time when the image was captured. The image recognition unit is configured to identify a set of images from the plurality of images, each image having a timestamp that is correlated with a timestamp associated with the objects and / or actions from a data stream. The data labeling unit is configured to determine, for each image in the set of images, an image label indicating a probability that the image depicts each type of object in a set of one or more specified object classes and / or a specified human action based on a correlation between the timestamp of the image and the timestamp associated with the objects and / or actions from the data stream.

[0006] Both Sharma and Ghafoor aim to improve image recognition methods by manipulating the training image sets used to train image recognition models. Thus, Sharma and Ghafoor's methods and systems increase the complexity of the already time-consuming task of developing image training sets.

[0007] Thus, there remains an opportunity to improve the accuracy of image recognition for part identification and assessment.Example systems and methods described herein are directed to overcoming one or more of the above-mentioned deficiencies. Summary of the Invention

[0008] In some embodiments, a method for object recognition using multiple images may include training an object recognition model by collecting multiple training images of each of a plurality of objects and labeling each of the multiple training images with a corresponding one of a plurality of object identifiers. Training the object recognition model may include training a neural network using the multiple labeled training images. The method may also include receiving at least two target images of a target object and feeding each of the at least two target images into the trained object recognition model to receive, for each of the at least two target images, a target object identifier corresponding to the target object and a probability that the target object identifier corresponds to the target object. The method may include calculating a similarity value between the at least two target images, combining the probabilities of the at least two target images in proportion to the similarity value, and identifying the target object as the target object identifier.

[0009] In some aspects, collecting the plurality of training images includes receiving a plurality of photographs of each of the plurality of objects. In some aspects, collecting the plurality of training images includes rendering a plurality of training images of each of the plurality of objects. In other aspects, rendering the plurality of training images includes rendering images of each of the plurality of objects viewed from different angles. In some aspects, the method further includes displaying information including an image of a part corresponding to the target object identifier located on an associated machine. In some aspects, the method further includes displaying a set of suitable alternative parts for the part corresponding to the target object identifier.

[0010] In some embodiments, an object recognition system may include one or more processors and one or more memory devices having instructions stored thereon. When executed by the one or more processors, the instructions cause the one or more processors to train an object recognition model by collecting a plurality of training images of each of a plurality of objects and labeling each of the plurality of training images with a corresponding one of a plurality of object identifiers. Training the object recognition model may also include training a neural network using the plurality of labeled training images. The instructions may also cause the one or more processors to receive at least two target images of a target object and feed each of the at least two target images into the trained object recognition model to receive, for each of the at least two target images, a probability that the target object corresponds to each of the plurality of object identifiers. The instructions may also cause the one or more processors to calculate a similarity value between the at least two target images and, for each of the plurality of object identifiers, combine the corresponding probabilities of the at least two target images in proportion to the similarity value. The instructions may also cause the one or more processors to identify the target object as the object identifier with the greatest combined probability.

[0011] According to some aspects, collecting a plurality of training images includes receiving a plurality of photographs of each of a plurality of objects. In some aspects, collecting a plurality of training images includes rendering a plurality of training images of each of the plurality of objects. In other aspects, rendering the plurality of training images includes rendering images of each of the plurality of objects viewed from different angles. In some aspects, the system further includes normalizing the combined probability of each object identifier. According to some aspects, the target object is a part associated with a machine, and the system further includes receiving machine information identifying the machine, and based on the machine information, removing the selected object identifier from the plurality of object identifiers. In some aspects, the system further includes displaying information including an image of the part corresponding to the object identifier with the greatest combined probability located on the associated machine.

[0012] In some embodiments, one or more non-transitory computer-readable media store computer-executable instructions that, when executed by one or more processors, cause the one or more processors to perform operations. The operations may include training an object recognition model by collecting a plurality of training images of each of a plurality of objects and labeling each of the plurality of training images with a corresponding one of a plurality of object identifiers. Training the wear estimation model may include training a neural network using the plurality of labeled training images. The operations may also include receiving at least two target images of a target object and feeding each of the at least two target images into the trained object recognition model to receive, for each of the at least two target images, a target object identifier corresponding to the target object and a probability that the target object identifier corresponds to the target object from the trained object recognition model. The operations may include calculating a similarity value between the at least two target images, combining the probabilities of the at least two target images in proportion to the similarity value, and identifying the target object as the target object identifier.

[0013] In some aspects, collecting the plurality of training images includes receiving a plurality of photographs of each of the plurality of objects. In other aspects, collecting the plurality of training images includes rendering a plurality of training images of each of the plurality of objects. In some aspects, rendering the plurality of training images includes rendering images of each of the plurality of objects viewed from different angles. According to some aspects, the operations may further include displaying information including an image of a part corresponding to the target object identifier located on an associated machine. In some aspects, the operations may further include displaying a set of suitable replacement parts for the part corresponding to the target object identifier. In some aspects, the operations may further include displaying ordering information for the part corresponding to the target object identifier. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The systems and methods described herein may be better understood by referring to the following detailed description in conjunction with the accompanying drawings, wherein like reference numerals indicate identical or functionally similar elements:

[0015] Figure 1 is a flow chart illustrating a method for object recognition using multiple images according to some embodiments of the disclosed technology;

[0016] Figure 2 is a flowchart illustrating a method for training an object recognition model according to some embodiments of the disclosed technology;

[0017] Figure 3A are the original photos of the parts used to train the object recognition model;

[0018] Figure 3B are augmented images used to train object recognition models;

[0019] Figure 4A is a perspective view depicting a rendered image of a machine part;

[0020] Figure 4B yes Figure 4A A side view of a rendered image of a machine part shown in ;

[0021] Figure 4C yes Figure 4A and Figure 4B A top view of a rendered image of a machine part shown in;

[0022] Figure 5 depicts renderings showing various image backgrounds and brightness levels used to create training images;

[0023] Figure 6 is a representative parts information and ordering display according to some embodiments of the disclosed technology;

[0024] Figure 7 is a block diagram showing an overview of a device on which some embodiments may operate;

[0025] Figure 8 is a block diagram illustrating an overview of an environment on which some embodiments may operate;

[0026] Figure 9 is a block diagram illustrating components that may be used in a system employing the disclosed technology in some embodiments.

[0027] The headings provided herein are for convenience only and do not necessarily affect the scope of the embodiments. In addition, the drawings are not necessarily rendered to scale. For example, the sizes of some elements in the drawings may be enlarged or reduced to help improve understanding of the embodiments. In addition, although the disclosed technology can be subjected to various modifications and alternative forms, specific embodiments are shown by way of example in the drawings and are described in detail below. However, the purpose is not to unnecessarily limit the embodiments described. On the contrary, the embodiments are intended to cover all modifications, combinations, equivalents and alternatives that fall within the scope of the present disclosure. DETAILED DESCRIPTION

[0028] Various examples of the systems and methods described above will now be described in greater detail. The following description provides specific details to facilitate a thorough understanding and implementation of the description of these examples. However, those skilled in the relevant art will appreciate that the techniques and technologies discussed herein can be practiced without many of these details. Similarly, those skilled in the relevant art will appreciate that the technology may include many other features not described in detail herein. Additionally, some well-known structures or functions may not be shown or described in detail below to avoid unnecessarily obscuring the relevant description.

[0029] The terms used below are to be interpreted in their broadest reasonable manner, even when used in conjunction with the detailed description of some specific examples of the embodiments. Indeed, some terms may even be emphasized below; however, any term intended to be interpreted in any restrictive manner will be explicitly and specifically defined as such in this section.

[0030] Methods and systems for object recognition using multiple images from different angles are disclosed. The techniques disclosed herein provide novel methods for combining probabilities derived from multiple images to improve part recognition accuracy. The disclosed techniques can also consider multiple images that are not all from different angles.

[0031] Identifying objects such as repair parts from images is often challenging due to the subtle differences in the parts, which require viewing the object from multiple angles. For example, to identify part X (e.g., a coupler with type A connector), one photo from above (top view) is needed to identify it as a coupler, and another photo from its axis down (front view) is needed to identify the connection type.

[0032] Typical image recognition algorithms (such as recurrent neural networks) operate on one image at a time. For example, from a top view, a typical algorithm can determine with 20% certainty that a part is part X. From a front view, the algorithm can determine with 50% certainty that the part is part X.

[0033] Assuming the images are independent, the probability that a part is part X can be calculated using the inclusion / exclusion rules for combining independent events, as follows:

[0034] 1-P(non-X|first image) x P(non-x|second image) = 1-0.8 x 0.5 = 0.6

[0035] Therefore, the probability that the part is part X using the combined image probability is 60%, which is higher than either individual evaluation using typical techniques. However, when the images are identical or very similar (e.g., from exactly the same angle), the above calculation provides erroneous results.

[0036] Recognizing an object from two images of completely different aspects of the object (e.g., from different angles) is an independent event. Recognizing an object from two very similar images is not independent, because the probability of recognition from one image is highly correlated with the probability of recognition from the other. In other words, multiple images of an object from the same angle contribute little additional information to the recognition process.

[0037] Assuming both photos are top views (with 20% certainty that the part is part X), the calculation is as follows:

[0038] 1-P(non-X|first image) x P(non-x|second image) = 1-0.8 x 0.8 = 0.36

[0039] Thus, using the inclusion / exclusion rule provides an erroneous 36% probability when the probability should be 20%. In this case, the additional photo does not provide any additional information on the initial 20% probability that the part is part X.

[0040] The disclosed technology can also combine recognition analysis based on multiple images in such a way that it considers when the additional images contribute additional information to the overall recognition (e.g., separate images at different angles) and when they do not contribute additional information (e.g., the same image). As explained more fully below, the similarity of the images is measured and the probability of combining the images is proportional to their similarity in order to consider similar images.

[0041] Figure 1 1 is a flow chart illustrating a method 100 for object recognition using multiple images according to some embodiments of the disclosed technology. The method 100 may include training an object recognition model at step 102, which will be discussed below with respect to Figure 2 As further described, an object recognition model is trained to recognize a plurality of object types, each object type being associated with an object identifier j.

[0042] At least two target images (i.e., image a and image b) of a target object are received at step 104. The target object may be a part of an asset such as a machine. The asset may include, for example, a truck, a crawler tractor, an excavator, a wheel loader, a front-end loader, and other equipment.

[0043] At step 106, each of the at least two target images is fed to a trained object recognition model, and for each of the at least two target images, a probability (e.g., p(a, j) and p(b, j)) that the target object corresponds to each of a plurality of object identifiers j is received.

[0044] At step 108, a similarity value s(a, b) is calculated between the two target images. In some embodiments, when there are more than two target images, the similarity is calculated for each pair of images. Any suitable similarity algorithm may be used to calculate the similarity between the two images. In some embodiments, a scale-invariant feature transform (SIFT) algorithm may be used to calculate the similarity. Regardless of the algorithm used, s(a, b) is scaled to a range of [0, 1], where 0 represents the same image and 1 represents completely different images.

[0045] For each of the plurality of object identifiers j, the corresponding probabilities of the target image (e.g., p(a, j) and p(b, j)) are combined in proportion to the similarity value (e.g., s(a, b)). Thus, at step 110, for each of the plurality of object identifiers j, the initial probability P(j) that the target image is of type j is set as follows, where n is the number of target object images:

[0046] P(j):=sum(a=1 to n){p(a,j)x[product(b=1 to a-1)(1-p(b,j))s(a,b)]}

[0047] At step 112, the initial probabilities P(j) are normalized to sum to 1 as follows, where k is the number of different object identifiers j:

[0048] P*(j):=P(j) / sum(a=1 to k)P(a).

[0049] At step 114, the target object is identified as the object identifier j having the largest combined probability. That is, the target object is identified as the object identifier x, where:

[0050] x=argmaxP*(j) where j=1...k.

[0051] Figure 2 is a flow chart illustrating a method 200 for training an object recognition model according to some embodiments of the disclosed technology. Method 200 may include collecting a plurality of training images of each of a plurality of objects (e.g., a library of training images). This may include collecting a plurality of training images in the form of photographs of each of the plurality of objects at step 202, and / or rendering a plurality of training images of each of the plurality of objects at step 204. At step 206, each of the plurality of training images (e.g., the photographs and the renderings) is labeled with a corresponding one of a plurality of object identifiers (e.g., a part number). Each image may also be labeled with a part name and various specifications, such as part weight, color, material, version, etc. At step 208, a neural network may be trained using the plurality of labeled training images. In some embodiments, the neural network may be a recurrent neural network.

[0052] The more and more diverse the training images in the library, the more likely an object recognition model (e.g., a neural network) is to correctly identify images in new situations, rather than overfitting on the limited training set. For example, if a model is trained only on images of clean parts that are well-lit, centered, and oriented at a 90-degree angle against a black background, it may have difficulty recognizing a dirty part that is off to the side in an image taken at an angle in poor lighting.

[0053] Figure 3A300 are original photos of parts used to train the object recognition model. Figure 3B Four augmented images 301-304 are shown that can also be used to train the model. Augmented images 301-304 are each rotated to provide different angles. These photos can be rotated using image editing software and / or taken from different angles. For more complex parts, different photos can be taken from different angles.

[0054] Generating a comprehensive library of training images from photographs can be time-consuming. However, virtual images can be rendered against a wide variety of backgrounds from any number of angles, under different lighting effects, with dirt and other obscuring materials. These images can be used in place of or in conjunction with actual photographs. Additionally, each correct recognition is added to the library of labeled training images.

[0055] The training images can be rendered with a computer-aided design (CAD) tool to create a photorealistic image of the part. This novel approach to developing a training library allows for the creation of many training images, each with precise and detailed labeling. Because the training images are generated from a CAD model, images of the part viewed from different angles can be easily created. Additionally, different lighting effects and backgrounds can be applied to the images. In some embodiments, simulated rust and dirt can be used to represent the typical condition of the part in use. Many CAD programs also allow for parametric modeling, by which the dimensions of a part can be managed by tabular data. For example, using parametric modeling, the creation of images of parts of different sizes and configurations can be automated. Additionally, parametric modeling can be used to automate the angles from which a part is viewed.

[0056] Figures 4A to 4C are various views depicting rendered images of a machine part 400 (eg, an excavator tooth) at different angles. Figure 4A A perspective view of an excavator tooth 400 is shown. Figure 4B A side view of an excavator tooth 400 is shown. Also, Figure 4C A top view of an excavator tooth 400 is shown. Each additional view (whether generated by rendering or photograph) adds additional information that is useful for training neural networks for image recognition. It should be understood that CAD programs have the ability to render photorealistic images that include surface textures, colors, and reflections representing different materials such as metal, plastic, and rubber, to name a few.

[0057] Figure 5Renderings are depicted showing various image backgrounds and brightness levels used to create training images of an excavator bucket. For example, in row 502, the image includes backgrounds ranging from a grassland, a wooded area, and a mountainous area. As shown in row 504, each of these rendered images can be adjusted by increasing the brightness of the image. As described above, with more variety in the training library, the object recognition model (e.g., a neural network) is more likely to correctly recognize the image in a new environment. As shown, the rendering can also include rust and dirt applied to the bucket and arm in a photorealistic manner.

[0058] To further promote accuracy, in some embodiments, a photograph of the target object can be taken from a standardized position with a standardized background, lighting, etc. The part can also be cleaned before photographing. In some embodiments, the photograph can be taken with a typical digital camera or cell phone camera. In some embodiments, the user can be prompted to take additional photographs of the target object until a threshold confidence level, such as 85%, is reached.

[0059] In some embodiments, limiting the scope of what object types should be considered can greatly increase the probability of correct identification of a part. For example, if the type and / or serial number of an asset is known, consideration of object types can be limited to only those parts that would fit that asset. Furthermore, if it is known that a specific object type (e.g., an excavator tooth) is being inspected, the scope can be narrowed down further. Thus, the user can be prompted to enter any known information about the associated machine and the part to be identified. In some embodiments, the user can capture a photo of a machine, and through image recognition of the machine, the user can be directed to the serial number location of that machine.

[0060] In some cases, two objects appear identical from every angle but have very different sizes, see e.g. Figure 3A Difficult to distinguish Figure 3A Different sizes of parts shown. As another example, consider two bolts, one of which is simply a magnified version of the other. To distinguish such objects, two images can be taken roughly simultaneously but from slightly different positions (e.g., images taken from the multiple lenses of a modern mobile phone), and the size of the target object can be estimated through stereo image disparity calculations.

[0061] Figure 66 is a representative part information and ordering display 600 according to some embodiments of the disclosed technology. Once a part is identified with sufficient confidence based on image recognition, information about the part can be displayed to the user, including an image 602 of the part located on an associated machine and / or subassemblies showing related parts. The user can also be prompted to verify that the part image 602 corresponds to the identified target object, thereby verifying that the object recognition model is correct. In some embodiments, the system can also display a set of suitable alternative parts. In some embodiments, the system can be configured to enable the part to be purchased (i.e., display an ordering page, add it to a shopping cart, or even automatically order it). The system can also display the specifications of the part (e.g., part weight, color, material, version, part number, etc.) and provide information related to installation procedures and maintenance.

[0062] In some embodiments, the above-described part identification technology can be employed to facilitate a system for identifying the condition of a part (e.g., dirty, rusted, bent, broken, or otherwise damaged or worn). By estimating the current wear on a part, the system can estimate the impact of cost on machine performance. Combined with information about machine usage, the system can determine when the cost of operating with a worn part will exceed the cost of replacing the part. This capability applies to parts that deteriorate due to wear, which can be visually detected as changes in the shape or appearance of the part. For example, but not limited to, such parts can include ground engaging tools, undercarriage parts, and tires. Assets include trucks, crawler tractors, excavators, wheel loaders, front-end loaders, and other equipment.

[0063] A method for optimizing part replacement time may include training a wear estimation model by predicting a plurality of wear patterns for a part, each wear pattern corresponding to a severity level; and rendering a plurality of training images, each training image representing a corresponding one of the plurality of wear patterns. Training the wear estimation model may include labeling each of the plurality of training images with a corresponding severity level, and training a neural network with the plurality of labeled training images. The method may also include receiving a part image of a deployed part associated with a machine, and feeding the part image into the trained wear estimation model to receive a wear estimate for the part image from the trained wear estimation model. The method may include estimating a change in machine performance based on the wear estimate and determining a machine usage pattern for the machine. The method may also include combining the machine usage pattern and the change in performance estimate to determine an optimal time to replace the part.

[0064] These methods and systems for part replacement time optimization are further described in co-pending U.S. application Ser. No. 17 / 123,058, filed on December 15, 2020, entitled “SYSTEMS AND METHODS FOR WEAR ASSESSMENT AND PART REPLACEMENT TIMING OPTIMIZATION” (Agent Docket No. 131257-8008.US00), the disclosure of which is incorporated herein by reference in its entirety.

[0065] Suitable system

[0066] The technology disclosed herein can be implemented as dedicated hardware (e.g., circuits), programmable circuits appropriately programmed with software and / or firmware, or a combination of dedicated and programmable circuits. Therefore, embodiments may include a machine-readable medium having instructions stored thereon, which instructions may be used to cause a computer, a microprocessor, a processor, and / or a microcontroller (or other electronic device) to perform processing. Machine-readable media may include, but are not limited to, optical disks, compact disk read-only memories (CD-ROMs), magneto-optical disks, ROMs, random access memories (RAMs), erasable programmable read-only memories (EPROMs), electrically erasable programmable read-only memories (EEPROMs), magnetic or optical cards, flash memories, or other types of media / machine-readable media suitable for storing electronic instructions.

[0067] Several embodiments are discussed in more detail below with reference to the accompanying drawings. Figure 7 7 is a block diagram illustrating an overview of a device on which some embodiments of the disclosed technology may operate. For example, a device may include hardware components of a device 700 that performs object recognition. The device 700 may include one or more input devices 720 that provide input to a CPU (processor) 710, thereby informing it of an action. The action is typically mediated by a hardware controller that interprets signals received from the input device and transmits the information to the CPU 710 using a communication protocol. The input device 720 includes, for example, a mouse, keyboard, touch screen, infrared sensor, touchpad, wearable input device, camera or image-based input device, microphone, or other user input device.

[0068] CPU 710 can be a single processing unit or multiple processing units in a device, or distributed across multiple devices. CPU 710 can be coupled to other hardware devices, for example, using a bus such as a PCI bus or a SCSI bus. CPU 710 can communicate with a hardware controller of a device, such as display 730. Display 730 can be used to display text and graphics. In some examples, display 730 provides graphical and textual visual feedback to the user. In some embodiments, display 730 includes an input device as part of the display, such as when the input device is a touch screen or is equipped with an eye direction monitoring system. In some embodiments, the display is separate from the input device. Examples of display devices include: LCD display screens; LED display screens; projection, holographic, or augmented reality displays (such as heads-up display devices or head-mounted devices); and the like. Other I / O devices 740 can also be coupled to the processor, such as a network card, video card, audio card, USB, FireWire, or other external devices, sensors, cameras, printers, speakers, CD-ROM drives, DVD drives, disk drives, or Blu-ray drives.

[0069] In some embodiments, the device 700 also includes a communication device capable of communicating with a network node wirelessly or by wire. The communication device can communicate with another device or server via a network using, for example, the TCP / IP protocol. The device 700 can utilize the communication device to distribute operations across multiple network devices.

[0070] CPU710 can access memory 750. Memory includes one or more of various hardware devices for volatile storage and non-volatile storage, and may include read-only memory and writable memory. For example, memory may include random access memory (RAM), CPU registers, read-only memory (ROM), and writable non-volatile memory, such as flash memory, hard drives, floppy disks, CDs, DVDs, magnetic storage devices, tape drives, device buffers, etc. Memory is not a propagating signal separate from the underlying hardware; therefore, memory is non-temporary. Memory 750 may include program memory 760 that stores programs and software, such as an operating system 762, an object recognition platform 764, and other applications 766. Memory 750 may also include data storage 770, which may include database information, etc., that may be provided to program memory 760 or any element of device 700.

[0071] Some embodiments can operate with numerous other general-purpose or special-purpose computing system environments or configurations. Examples of well-known computing systems, environments, and / or configurations suitable for use with the technology include, but are not limited to, personal computers, server computers, handheld or laptop devices, cellular phones, mobile phones, wearable electronic devices, game consoles, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like.

[0072] Figure 8 8 is a block diagram illustrating an overview of an environment 800 in which some embodiments of the disclosed technology may operate. The environment 800 may include one or more client computing devices 805A-D, instances of which may include the device 700. The client computing device 805 may operate in a network environment using a logical connection to one or more remote computers, such as a server computing device 810, over a network 830.

[0073] In some embodiments, server computing device 810 may be an edge server that receives client requests and coordinates the fulfillment of these requests through other servers such as servers 820A-C. Server computing devices 810 and 820 may include computing systems such as device 700. Although each server computing device 810 and 820 is logically shown as a single server, each server computing device may be a distributed computing environment including multiple computing devices located at the same or geographically different physical locations. In some embodiments, each server computing device 820 corresponds to a group of servers.

[0074] The client computing device 805 and the server computing devices 810 and 820 can each act as a server or client to the other server / client devices. Server 810 can be connected to a database 815. Servers 820A-C can each be connected to a corresponding database 825A-C. As described above, each server 820 can correspond to a group of servers, and each of these servers can share a database or can have its own database. Databases 815 and 825 can store (e.g., store) information. Although databases 815 and 825 are logically shown as a single unit, databases 815 and 825 can each be a distributed computing environment including multiple computing devices, which can be located within their corresponding servers, or can be located at the same or geographically different physical locations.

[0075] The network 830 can be a local area network (LAN) or a wide area network (WAN), but can also be other wired or wireless networks. The network 830 can be the Internet or some other public or private network. The client computing device 805 can be connected to the network 830 via a network interface (e.g., via wired or wireless communication). Although the connection between the server 810 and the server 820 is shown as a separate connection, these connections can be any type of local, wide area, wired or wireless network, including the network 830 or a separate public or private network.

[0076] Figure 9 9 is a block diagram illustrating components 900 that may be used in a system employing the disclosed technology in some embodiments. Components 900 include hardware 902, general-purpose software 920, and specialized components 940. As described above, a system implementing the disclosed technology may use various hardware, including a processing unit 904 (e.g., a CPU, GPU, APU, etc.), a working memory 906, a storage memory 908, and input and output devices 910. Components 900 may be implemented in a client computing device such as client computing device 805 or on a server computing device such as server computing devices 810 or 820.

[0077] General software 920 may include various application programs, including an operating system 922, local programs 924, and a basic input and output system (BIOS) 926. Specialized components 940 may be subcomponents of general software application programs 920, such as local programs 924. Specialized components 940 may include a model training module 944, an object rendering module 946, an object recognition module 948, a parts information and ordering module 950, and components that may be used to transfer data and control specialized components, such as an interface 942. In some embodiments, component 900 may be in a computing system distributed across multiple computing devices, or may be an interface to a server-based application that executes one or more specialized components 940.

[0078] Those skilled in the art will understand that the above Figures 7 to 9 The components shown in each of the above flow charts can be modified in various ways. For example, the order of the logic can be rearranged, sub-steps can be performed in parallel, the logic shown can be omitted, other logic can be included, etc. In some embodiments, one or more of the above components can perform one or more processes described herein.

[0079] Industrial Applicability

[0080] In some embodiments, an object recognition system using multiple images may include a model training module 944, an object rendering module 946, an object recognition module 948, and a parts information and ordering module 950 ( Figure 9In operation, the model training module 944 may train the object recognition model by collecting a plurality of training images of each of a plurality of objects and labeling each of the plurality of training images with a corresponding one of a plurality of object identifiers. Training the object recognition model may also include training a neural network using the plurality of labeled training images. In some embodiments, the collected training images may include photographs and rendered training images of each of the plurality of objects generated by the object rendering module 946. The rendered training images may include rendered images of each of the plurality of objects viewed from different angles.

[0081] The object recognition module 948 may receive at least two target images of a target object and feed each of the at least two target images into a trained object recognition model to receive, from the trained object recognition model, a probability that the target object corresponds to each of a plurality of object identifiers for each of the at least two target images. The object recognition module 948 may also calculate a similarity value between the at least two target images and, for each of the plurality of object identifiers, combine the corresponding probabilities of the at least two target images in proportion to the similarity value. The object recognition module 948 causes the one or more processors to identify the target object as the object identifier having the greatest combined probability.

[0082] Once a part has been identified with sufficient confidence (e.g., greater than a threshold probability) based on image recognition, information about the part can be displayed to the user via the part information and ordering module 950. This information can include an image of the part located on the associated machine and / or subassemblies showing the related parts. In some embodiments, the system can also display a set of suitable replacement parts. The system can be configured to enable the part to be purchased (i.e., display an ordering page, add it to a shopping cart, or even automatically order it), display the specifications of the part (e.g., part weight, color, material, version, part number, etc.) and provide information related to installation procedures and maintenance. In some embodiments, the above-described part recognition technology can be employed to facilitate a system for identifying part conditions (e.g., dirty, rusted, bent, broken, or otherwise damaged or worn) and optimizing part replacement time.

[0083] Remark

[0084] The above description and accompanying drawings are illustrative and should not be construed as restrictive. Many specific details are described to provide a thorough understanding of the present disclosure. However, in some cases, well-known details are not described to avoid obscuring the description. In addition, various changes may be made without departing from the scope of the embodiments.

[0085] References in this specification to "one embodiment" or "an embodiment" mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present disclosure. The appearances of the phrase "in one embodiment" in various places in the specification are not necessarily all referring to the same embodiment, nor are separate or alternative embodiments mutually exclusive of other embodiments. Furthermore, various features are described that may be exhibited by some embodiments but not others. Similarly, various requirements are described that may be requirements for some embodiments but not others.

[0086] The terms used in this specification generally have their ordinary meaning in the art, in the context of the present disclosure, and in the specific context of using each term. It should be understood that the same thing can be explained in more than one way. Therefore, alternative language and synonyms can be used for any one or more of the terms discussed herein, and no matter whether the terms are elaborated or discussed in this article, no special meaning is set for them. Synonyms for some terms are provided. The description of one or more synonyms does not exclude the use of other synonyms. The use of examples anywhere in this specification (including examples of any terms discussed herein) is only illustrative and is not intended to further limit the scope and meaning of the present disclosure or any exemplary term. Similarly, the present disclosure is not limited to the various embodiments provided in this specification. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those of ordinary skill in the art to which the present disclosure belongs. In the event of conflict, this document (including definitions) shall prevail.

Claims

1. A method (100) for object recognition using multiple images, comprising: Training an object recognition model (200) includes: collecting a plurality of training images of each of a plurality of objects (202); labeling each of the plurality of training images with a corresponding one of a plurality of object identifiers (206); and training a neural network using a plurality of labeled training images (208); receiving at least two target images (104) of a target object; feeding each of the at least two target images into a trained object recognition model (106); receiving, for each of the at least two target images, a target object identifier corresponding to the target object and a probability that the target object identifier corresponds to the target object from the trained object recognition model (106); Calculating a similarity value between the at least two target images (108); combining the probabilities of the at least two target images in proportion to the similarity values ​​(110); and The target object is identified as the target object identifier having the greatest combined probability (114).

2. The method of claim 1, wherein collecting the plurality of training images comprises receiving a plurality of photographs of each of the plurality of objects (202).

3. The method of claim 1, wherein collecting the plurality of training images comprises rendering a plurality of training images of each of the plurality of objects (204).

4. An object recognition system (700), comprising: one or more processors (710); as well as One or more memory devices (760) having stored thereon instructions that, when executed by the one or more processors (710), cause the one or more processors to: Training an object recognition model (200) includes: collecting a plurality of training images of each of a plurality of objects (202); labeling each of the plurality of training images with a corresponding one of a plurality of object identifiers (206); and training a neural network using a plurality of labeled training images (208); receiving at least two target images (104) of a target object; feeding each of the at least two target images into a trained object recognition model (106); receiving, for each of the at least two target images, from the trained object recognition model a probability that the target object corresponds to each of the plurality of object identifiers (106); Calculating a similarity value between the at least two target images (108); combining, for each of the plurality of object identifiers, corresponding probabilities of the at least two target images in proportion to the similarity value (110); and The target object is identified as the object identifier having the greatest combined probability (114).

5. The system of claim 4, wherein collecting the plurality of training images comprises receiving a plurality of photographs of each of the plurality of objects (202).

6. The system of claim 4, wherein collecting the plurality of training images comprises rendering a plurality of training images of each of the plurality of objects (204).

7. The system of claim 6, wherein rendering the plurality of training images comprises rendering an image of each of the plurality of objects viewed from a different angle (204).

8. One or more non-transitory computer-readable media (760) storing computer-executable instructions that, when executed by one or more processors, cause the one or more processors to perform operations including: Training an object recognition model (200) includes: collecting a plurality of training images of each of a plurality of objects (202); labeling each of the plurality of training images with a corresponding one of a plurality of object identifiers (206); as well as training a neural network using a plurality of labeled training images (208); receiving at least two target images (104) of a target object; feeding each of the at least two target images into a trained object recognition model (106); receiving, for each of the at least two target images, a target object identifier corresponding to the target object and a probability that the target object identifier corresponds to the target object from the trained object recognition model (106); Calculating a similarity value between the at least two target images (108); combining the probabilities of the at least two target images in proportion to the similarity values ​​(110); and The target object is identified as the target object identifier having the greatest combined probability (114).

9. The one or more non-transitory computer-readable media of claim 8, wherein collecting the plurality of training images comprises rendering a plurality of training images of each of the plurality of objects (204).

10. The one or more non-transitory computer-readable media of claim 9, wherein rendering the plurality of training images comprises rendering an image of each of the plurality of objects viewed from a different angle (204).

Citation Information

Patent Citations

  • Generating labeled training images for use in training a computational neural network for object or action recognition

    US20190392318A1

  • Systems and methods for wear assessment and part replacement timing optimization

    US20220188774A1

  • Image processing arrangements

    US10664722B1

  • System and method for enabling search and retrieval from image files based on recognized information

    US20060253491A1