Image enhancement techniques for automatic visual inspection
Image augmentation techniques using arithmetic transposition and inpainting generate causally relevant synthetic images, improving AVI system performance and reducing the complexity and cost of library development.
Patent Information
- Application Number
- JP2025161305
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-12-02
- Filing Date
- 2025-09-29
- Publication Date
- 2026-01-06
AI Technical Summary
Existing automated visual inspection (AVI) systems, particularly those using deep learning, require large and diverse image sets to develop and train models effectively, but current image libraries can lead to poor performance due to focusing on irrelevant features, and changes in products or processes necessitate iterative and costly rebuilding of these libraries.
Implementing image augmentation techniques such as arithmetic transposition and digital inpainting to generate causally relevant synthetic images that expand and balance the training library, ensuring pixel-level realism and reducing reliance on background bias.
Enhances AVI performance by minimizing false accepts and false rejects, reducing the complexity and cost of developing image libraries, and facilitating rapid adaptation to product or process changes.
Smart Images

Figure 2026001115000001_ABST
Abstract
Description
[Technical Field]
[0001] This application relates generally to automated visual inspection systems for pharmaceutical or other applications, and more particularly to techniques for expanding image libraries for use in developing, training, and / or validating such systems. [Background technology]
[0002] In various contexts, quality control procedures require samples to be carefully inspected for defects, and any samples that exhibit defects are rejected, discarded, and / or further analyzed. For example, in the pharmaceutical manufacturing context, containers (e.g., syringes or vials) and / or their contents (e.g., liquid or lyophilized pharmaceutical products) must be rigorously inspected for defects before sale or distribution. Similarly, many other industries rely on visual inspection to ensure product quality or for other purposes. To eliminate human error, reduce costs, and / or shorten inspection time (e.g., due to the large volumes of drugs or other items in commercial production), defect inspection tasks are increasingly being automated (i.e., "automated visual inspection" or "AVI"). For example, "computer vision" or "machine vision" software has been used in the pharmaceutical context.
[0003] Recently, deep learning techniques have emerged as a promising tool for AVI. However, these techniques generally require significantly more images than traditional AVI systems to develop, train, and thoroughly test models (e.g., neural networks). Furthermore, robust model performance generally relies on carefully designed image sets. For example, the image sets should represent a sufficient variety of conditions (e.g., by showing defects in different locations and having a range of different shapes and sizes). Furthermore, even with a large and diverse training image library, poor AVI performance can result if the image set causes the deep learning model to make decisions for the wrong reasons (e.g., based on irrelevant image features). This can be particularly problematic in contexts or scenarios where the defects shown are small or inoffensive relative to other (defect-free) image features.
[0004] For both deep learning and more traditional (e.g., machine vision) AVI systems, the development and qualification process using sample image libraries should ensure that false negatives or “false accepts” (i.e., missing defects) and false positives or “false rejects” (i.e., incorrectly identifying defects) are within acceptable thresholds. For example, in certain contexts (e.g., pharmaceutical contexts where patient safety is a concern), zero or near-zero false negatives may be required. While false positives may not be as critical, they can be very costly in economic terms and more difficult to address than false negatives when developing an AVI system. These and other factors can make image library development a highly iterative process that is very complex, labor-intensive, and costly. Furthermore, changes in the product line (e.g., new drugs, new containers, new fill levels for drugs within containers, etc.) or changes in the inspection process itself (e.g., different types of camera lenses, changes in camera position or lighting, etc.) can necessitate not only retraining and / or requalification of the model, but also (potentially) partial or complete rebuilding of the image library. Summary of the Invention [Means for solving the problem]
[0005] Embodiments described herein relate to automated image augmentation techniques that aid in the generation and / or evaluation of image libraries for developing, training, and / or validating robust deep learning models for AVI. In particular, various image augmentation techniques disclosed herein apply digital transformations to “original” images to artificially expand the scope of a training library (e.g., for deep learning AVI applications or more traditional computer / machine vision AVI applications). Unlike relatively simple image transformations (e.g., reflection, linear scaling, and rotation) previously used to expand image libraries, the techniques described herein can facilitate the generation of libraries that are not only larger and more diverse, but also more balanced and “causally relevant,” i.e., more likely to make classifications / decisions for the right reasons rather than focusing on irrelevant image features, and therefore more likely to provide good performance across a wide range of samples. To ensure causality, implementations described herein are used to generate a large number of “population-representative” synthetic images (i.e., synthetic images that are sufficiently representative of the images inferred by the model during runtime operation).
[0006] In one aspect of the present disclosure, a novel arithmetic transposition algorithm is used to generate a synthetic image from an original image by transposing features onto the original image with pixel-level realism. The arithmetic transposition algorithm can be used to generate a synthetic "flawed" image (i.e., an image that shows defects) by augmenting a "good" image (i.e., an image that does not show those defects) with images of the defects themselves. As an example, the algorithm can use images of a defect-free syringe and images of the syringe to generate a synthetic image of a syringe with cracks, deformed plungers, and / or other defects. As another example, the algorithm can use images of a defect-free body part and images of the defects to generate a synthetic image of an automobile body part with chips, scratches, dents, and / or other defects. Numerous other applications are also possible in quality control or other contexts.
[0007] In other aspects of the present disclosure, digital “inpainting” techniques are used to generate realistic synthetic images from original images to supplement image libraries for training and / or validation of AVI models (e.g., deep learning-based AVI models). In one such aspect, defects shown in the original image can be removed by masking the defects in the original image, calculating correspondence metrics between (1) portions of the original image adjacent to the masked region and (2) other portions of the original image outside the masked region, and filling the masked region with artificial, defect-free portions based on the calculated metrics. The ability to remove defects from an image can have a subtle yet significant impact on the training image library. In particular, complementary “good” and “faulty” images can be used in tandem to minimize the effects of background bias when training an AVI model.
[0008] Other digital inpainting techniques of the present disclosure leverage deep learning, such as deep learning based on partial convolution. Variations of these deep learning-based inpainting techniques can be used to remove defects from an original image, add defects to an original image, and / or modify (e.g., move or change the appearance of) features in an original image. For example, variations of these techniques can be used to remove cracks, chips, fibers, deformed plungers, or other defects from an image of a syringe containing pharmaceutical product, add such defects to a syringe image where the defects were not originally shown, or move or otherwise modify a meniscus or plunger shown in the original syringe image. These deep learning-based inpainting techniques facilitate careful design of training image libraries and can provide a good solution even for high-mix, low-volume production applications where it has traditionally been difficult to develop training image libraries in a cost-effective manner.
[0009] In general, the image enhancement techniques disclosed herein can improve AVI performance for both "false accepts" and "false rejects." Image enhancement techniques that add variability to the presented attributes / features (e.g., meniscus level, void size, air bubbles, small irregularities in glass container walls, etc.) can be particularly useful for reducing false rejects.
[0010] In yet other aspects of the present disclosure, quality control techniques are used to evaluate the suitability of image libraries for training and / or validating AVI deep learning models and / or to evaluate the suitability of individual images for inclusion in such libraries. These can include both "pre-processing" quality control techniques that evaluate image variability across a dataset, and "post-processing" quality control techniques that evaluate the degree of similarity between synthetic / augmented images and a set of images (e.g., real images that have not been altered by adding, removing, or modifying indicated features).
[0011] Those skilled in the art will understand that the figures described herein are included for illustrative purposes and are not intended to limit the present disclosure. The drawings are not necessarily to scale, emphasis instead being placed upon illustrating the principles of the present disclosure. It should be understood that in some instances, various aspects of the described implementations may be shown exaggerated or enlarged to facilitate understanding of the described implementations. [Brief explanation of the drawings]
[0012] [Figure 1] FIG. 1 is a simplified block diagram of an example system in which various techniques described herein relating to the development and / or evaluation of an Automated Visual Inspection (AVI) image library may be implemented. [Figure 2] FIG. 2 illustrates an exemplary visual inspection system that may be used in a system such as the system of FIG. 1. [Figure 3A] 3A-3C illustrate various exemplary container types that may be inspected using a visual inspection system such as the visual inspection system of FIG. 2. [Figure 3B] 3A-3C illustrate various exemplary container types that may be inspected using a visual inspection system such as the visual inspection system of FIG. 2. [Figure 3C] 3A-3C illustrate various exemplary container types that may be inspected using a visual inspection system such as the visual inspection system of FIG. 2. [Figure 4A-1] FIG. 1 illustrates an arithmetic transposition algorithm that can be used to add features to images with pixel-level realism. [Figure 4A-2] FIG. 1 illustrates an arithmetic transposition algorithm that can be used to add features to images with pixel-level realism. [Figure 4B] FIG. 4B illustrates an exemplary defect matrix histogram that may be generated during the arithmetic transpose algorithm of FIG. 4A. [Figure 4C] FIG. 4B illustrates an exemplary defect matrix histogram that may be generated during the arithmetic transpose algorithm of FIG. 4A. [Figure 5] FIG. 10 is a diagram illustrating an example of an operation in which a feature image is converted into a numerical matrix. [Figure 6] FIG. 6 compares a manually generated image of a syringe with a real-world crack with a synthetic image of a syringe with a digitally generated crack, the synthetic image being generated using the arithmetic transposition algorithm of FIG. 5. [Figure 7] FIG. 8 is a comparison of pixel levels corresponding to the images of FIG. 7. [Figure 8] FIG. 6 compares defects synthesized using the prior art with defects synthesized using the arithmetic transpose algorithm of FIG. 5. [Figure 9A] 6A-6C show various composite images with defects produced using the arithmetic transposition algorithm of FIG. 5. [Figure 9B] FIG. 6 illustrates a collection of exemplary crack defect images, each of which may be used as input to the arithmetic transposition algorithm of FIG. 5. [Figure 10] FIG. 10 illustrates a heat map used to evaluate the effectiveness of the augmented image. [Figure 11] 10 is a plot showing AVI neural network performance for different combinations of synthetic and real images in training and test image sets. [Figure 12] FIG. 1 illustrates an exemplary partial convolution model that can be used to generate a synthetic image by adding, removing, or modifying the indicated features. [Figure 13] FIG. 10 illustrates example masks that can be randomly generated for use in training a partially convolutional model. [Figure 14-1] 1A-1C show three example sequences in which synthetic images are generated by digitally removing defects from real images using a partial convolution model. [Figure 14-2] 1A-1C show three example sequences in which synthetic images are generated by digitally removing defects from real images using a partial convolution model. [Figure 15]FIG. 10 shows another example of a synthetic image produced by digitally removing defects from a real image using a partial convolution model, accompanied by a difference image showing how the real image has been modified. [Figure 16] FIG. 1 shows a real image of a defective syringe and a composite image of a non-defective syringe, where the composite image is generated based on the real image using a partial convolution model. [Figure 17] 1A-1C show three exemplary defect images that can be used with a partial convolution model to digitally add defects to a syringe image according to a first technique. [Figure 18] 1A-1C illustrate two exemplary sequences in which a partial convolution model is used to add defects to a syringe image according to a first technique. [Figure 19] FIG. 1 shows a real image of a syringe without defects and a composite image of a syringe with defects, where the composite image is generated based on the real image using a partial convolution model and a first technique. [Figure 20-1] 10A-10C show three exemplary sequences in which a partial convolution model is used to add defects to syringe images according to a second technique. [Figure 20-2] 10A-10C show three exemplary sequences in which a partial convolution model is used to add defects to syringe images according to a second technique. [Figure 21] FIG. 1 shows a real image of a syringe without defects and a synthetic image of a syringe with defects, where the synthetic image is generated based on the real image using a partial convolution model and a second technique. [Figure 22] 10A-10C illustrate an exemplary sequence in which a partial convolution model is used to correct a meniscus in a syringe image according to a second technique. [Figure 23] FIG. 1 shows a real image of a syringe and a composite image in which the meniscus has been digitally altered, the composite image being generated based on the real image using a partial convolution model and a second technique. [Figure 24A]FIG. 10 illustrates exemplary heatmaps showing the causal relationships underlying predictions made by AVI deep learning models trained with and without synthetic training images. [Figure 24B] FIG. 10 illustrates exemplary heatmaps showing the causal relationships underlying predictions made by AVI deep learning models trained with and without synthetic training images. [Figure 25] FIG. 1 illustrates an exemplary process for generating a visualization that can be used to assess diversity in a set of images. [Figure 26A] FIG. 26 illustrates an exemplary visualization produced by the process of FIG. 25. [Figure 26B] FIG. 10 illustrates an exemplary visualization that can be used to assess the diversity of an image set using another process. [Figure 27] FIG. 1 illustrates an exemplary process for assessing similarity between a composite image and a set of images. [Figure 28] FIG. 28 illustrates an exemplary histogram generated using the process of FIG. 27. [Figure 29] FIG. 1 is a flow diagram of an exemplary method for generating a synthetic image by transferring features onto an original image. [Figure 30] FIG. 1 is a flow diagram of an exemplary method for generating a composite image by removing defects shown in an original image. [Figure 31] FIG. 1 is a flow diagram of an exemplary method for generating a composite image by removing or modifying features shown in the original image or by adding features shown in the original image. [Figure 32] FIG. 1 is a flow diagram of an exemplary method for evaluating synthetic images used in a training image library. DETAILED DESCRIPTION OF THE INVENTION
[0013] The various concepts described introductory above and discussed in more detail below can be implemented in any of many ways, and the concepts described are not limited to any particular mode of implementation. Example implementations are provided for illustrative purposes.
[0014] As used herein (and used interchangeably), the terms "synthetic image" and "augmented image" generally refer to an image that has been digitally altered to show something different from what the image originally showed, and should be distinguished from the output produced by other types of image processing (e.g., adjusting contrast, changing resolution, cropping, filtering, etc.) that do not change the essence of what is shown. Conversely, a "real image" as referred to herein refers to an image that is not a synthetic / augmented image, regardless of whether other types of image processing have previously been applied to the image. An "original image" as referred to herein is an image that is being digitally modified to generate a synthetic / augmented image, and may be a real image or a synthetic image (e.g., an image that has been previously augmented, prior to an additional round of augmentation). References herein to "features" (e.g., "defects") are references to features of what is imaged (e.g., a crack or meniscus on a syringe shown in an image of a syringe, or a scratch or dent on an automobile body part shown in an image of a component), and are distinguished from features of the image itself that are unrelated to the nature of what is imaged (e.g., missing or damaged portions of an image, such as faded or dirty portions of an image).
[0015] 1 is a simplified block diagram of an example system 100 in which various techniques described herein related to the development and / or evaluation of an automated visual inspection (AVI) training and / or validation image library may be implemented. For example, the image library may be used to train one or more neural networks to perform AVI tasks. Once trained and qualified, the AVI neural network may be used for quality control during manufacturing (and / or in other contexts) to detect defects. In a pharmaceutical context, for example, an AVI neural network may be used to detect defects associated with syringes, vials, cartridges, or other container types (e.g., cracks, scratches, stains, missing components, etc.) and / or defects associated with fluids or lyophilized pharmaceutical products within the container (e.g., the presence of fibers and / or other foreign particles). As another example, in the automotive context, AVI neural networks may be used to detect defects (e.g., cracks, scratches, dents, stains, etc.) in the body structure of an automobile or other vehicle during manufacturing and / or at other times (e.g., checking the condition of returned rental vehicles to help determine fair resale value). Numerous other uses are also possible. Because the disclosed techniques can substantially reduce the cost and time associated with building image libraries, AVI neural networks may be used to detect visible defects in virtually any quality control application (e.g., checking the condition of home appliances, house siding, textiles, glassware, etc. before sale). While the examples provided herein relate primarily to the medical context, it will be understood that the techniques described herein need not be limited to such applications. Furthermore, in some implementations, synthetic images are used for purposes other than training an AVI neural network. For example, the images may instead be used to qualify systems that use computer vision without using deep learning.
[0016] System 100 includes a visual inspection system 102 configured to generate training and / or validation images. Specifically, visual inspection system 102 includes hardware (e.g., a transport mechanism, a light source, a camera, etc.) and firmware and / or software configured to capture digital images of a sample (e.g., a container holding a fluid or a lyophilized substance). An example of visual inspection system 102 is described below with reference to FIG. 2 , although any suitable visual inspection system may be used. In some embodiments, visual inspection system 102 is an offline (e.g., lab-based) “mock station” that closely replicates key aspects (e.g., optics, lighting, etc.) of a commercial line equipment station, thereby enabling the development of training and / or validation libraries without causing excessive downtime of the commercial line equipment. The development, deployment, and use of exemplary imitation stations are shown and discussed in PCT Patent Application No. PCT / US20 / 59776, entitled "Offline Troubleshooting and Development for Automated Visual Inspection Stations," filed November 10, 2020, which is incorporated herein by reference in its entirety. In other embodiments, the visual inspection system 102 is a commercial line of equipment that is also used during production.
[0017] The visual inspection system 102 may sequentially image each of a plurality of samples (e.g., containers). To this end, the visual inspection system 102 may include or operate in conjunction with a Cartesian robot, conveyor belt, carousel, star wheel, and / or other transport means that sequentially moves each sample to an appropriate position for imaging and then moves the sample away once imaging of the sample is complete. Although not shown in FIG. 1 , the visual inspection system 102 may include a communications interface and processor that enables communication with the computer system 104.
[0018] The computer system 104 may generally be configured to control / automate the operation of the visual inspection system 102 and to receive and process images captured / generated by the visual inspection system 102, as described further below. The computer system 104 may be a general-purpose computer specially programmed to perform the operations discussed herein, or may be a dedicated computing device. As seen in FIG. 1 , the computer system 104 includes a processing unit 110 and a memory unit 114. However, in some embodiments, the computer system 104 includes two or more computers that are co-located with each other or remote from each other. In these distributed embodiments, the operations described herein related to the processing unit 110 and the memory unit 114, or related to any of the modules performed when the processing unit 110 executes instructions stored in the memory unit 114, may be divided among multiple processing units and / or multiple memory units.
[0019] Processing unit 110 includes one or more processors, each of which may be a programmable microprocessor that executes software instructions stored in memory unit 114 to perform some or all of the functionality of computer system 104 described herein. Processing unit 110 may include, for example, one or more graphics processing units (GPUs) and / or one or more central processing units (CPUs). Alternatively, or in addition, one or more processors in processing unit 110 may be other types of processors (e.g., application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), etc.), and some of the functionality of computer system 104 described herein may instead be implemented in hardware.
[0020] Memory unit 114 may include one or more volatile and / or non-volatile memories. Any suitable memory type may be included in memory unit 114, such as read-only memory (ROM) and / or random access memory (RAM), flash memory, solid-state drive (SSD), hard disk drive (HDD), etc. Collectively, memory unit 114 may store one or more software applications, data received / used by those applications, and data output / generated by those applications.
[0021] In particular, memory unit 114 stores software instructions for various modules that, when executed by processing unit 110, perform various functions for the purposes of training, validating, and / or qualifying one or more AVI neural networks and / or other types of AVI software (e.g., computer vision software). Specifically, in the exemplary embodiment of FIG. 1, memory unit 114 includes an AVI neural network module 120, a visual inspection system (VIS) control module 122, a library expansion module 124, and an image / library evaluation module 126. In other embodiments, memory unit 114 may omit one or more of modules 120, 122, 124, and 126 and / or may include one or more additional modules. As noted above, computer system 104 may be a distributed system, in which case one, some, or all of modules 120, 122, 124, and 126 may be implemented in whole or in part by different computing devices or systems (e.g., by a remote server coupled to computer system 104 via one or more wired and / or wireless communication networks). Additionally, the functionality of any one of modules 120, 122, 124, and 126 may be divided among different software applications. By way of example only, in an embodiment in which computer system 104 accesses a web service to train and use one or more AVI neural networks, some or all of the software instructions of AVI neural network module 120 may be stored and executed on a remote server.
[0022] AVI neural network module 120 includes software that trains one or more AVI neural networks using images stored in training image library 140. Training image library 140 may be stored in memory unit 114 and / or in another local or remote memory (e.g., memory coupled to a remote library server, etc.). In some embodiments, in addition to training, AVI neural network module 120 can implement / execute the trained AVI neural network by, for example, applying newly acquired images by visual inspection system 102 (or another visual inspection system) to the neural network for validation, qualification, or possibly runtime operation. In various embodiments, the AVI neural network trained by AVI neural network module 120 classifies the image as a whole (e.g., defect vs. no defect, or the presence or absence of a particular type of defect), classifies the image pixel by pixel (i.e., image segmentation), detects objects within the image (e.g., detects the presence and location of a particular defect type, such as a scratch, crack, or foreign object), or a combination thereof (e.g., one neural network classifies the image and another neural network performs object detection). In some implementations, AVI neural network module 120 generates a heat map related to the operation of the trained AVI neural network (for reasons described below). To this end, AVI neural network module 120 may include deep learning software such as MVTec from HALCON®, Vidi® from Cognex®, Rekognition® from Amazon®, TensorFlow, PyTorch, and / or other suitable off-the-shelf or customized deep learning software. The software of the AVI neural network module 120 may be built on one or more pre-trained networks, such as, for example, ResNet50 or VGGNet, and / or one or more custom networks.
[0023] In some embodiments, the VIS control module 122 controls / automates the operation of the visual inspection system 102 so that sample images (e.g., container images) can be generated with little or no human interaction. The VIS control module 122 may cause a given camera to capture a sample image by sending a command or other electronic signal to the camera (e.g., generating a pulse on a control line). The visual inspection system 102 may send the captured container image to the computer system 104, which can store the image in the memory unit 114 for local processing. In alternative embodiments, the visual inspection system 102 may be controlled locally, in which case the VIS control module 122 may have fewer functions than those described herein (e.g., only handle the retrieval of images from the visual inspection system 102) or may be omitted from the memory unit 114 entirely.
[0024] The library augmentation module 124 (also referred to herein simply as “module 124”) processes sample images generated by the visual inspection system 102 (and / or other visual inspection systems) to generate additional, synthesized / augmented images for inclusion in the training image library 140. Module 124 can implement one or more image augmentation techniques, including any one or more of the image augmentation techniques disclosed herein. As described below, some of these image augmentation techniques can utilize the feature image library 142 to generate the synthesized images. The feature image library 142 may be stored in the memory unit 114 and / or another local or remote memory (e.g., memory coupled to a remote library server, etc.) and includes images of various types of defects (e.g., cracks, scratches, chips, stains, foreign objects, etc.) and / or images of variations of each defect type (e.g., cracks of different sizes and / or patterns, foreign objects with different shapes and sizes, etc.). Alternatively, or additionally, the feature image library 142 may include images of various other types of features (e.g., different menisci) that may or may not be indicative of defects. The images in feature image library 142 may be, for example, cropped portions of full sample images such that a substantial portion of each image contains a feature (eg, a defect).
[0025] In general, feature image library 142 can include images of virtually any type of feature associated with the sample being imaged. For example, in a pharmaceutical context, feature image library 142 may include defects associated with containers (e.g., syringes, cartridges, vials, etc.), container contents (e.g., liquid or lyophilized pharmaceutical products), and / or interactions between the containers and their contents (e.g., leaks, etc.). By way of non-limiting example, defect images may include images of syringe defects such as cracks, chips, scratches, and / or scrapes in the barrel, shoulder, neck, or flange, broken or deformed flanges, air lines in the glass of the barrel, shoulder, or neck wall, discontinuities in the glass of the barrel, shoulder, or neck, dirt on the inside or outside (or inside) of the barrel, shoulder, or neck wall, glass buildup on the barrel, shoulder, or neck, knots in the barrel, shoulder, or neck wall, foreign particles embedded within the glass of the barrel, shoulder, or neck wall, a foreign, misaligned, missing, or extra plunger, dirt on the plunger, deformed ribs on the plunger, incomplete or peeling coating on the plunger, a plunger in an unacceptably positioned position, a missing, bent, deformed, or broken needle shield, a needle protruding through the needle shield, etc. Examples of defects related to the interaction between the syringe and its contents may include fluid leaking through the plunger, fluid in the ribs of the plunger, fluid leaking through the needle shield, etc. The various components of an exemplary syringe are shown in FIG. 3A, discussed below.
[0026] Non-limiting examples of defects related to the cartridge may include cracks, chips, scratches, and / or scrapes on the barrel or flange, broken or deformed flanges, discontinuities in the barrel, dirt on the inside or outside (or inside) of the barrel, material stuck to the barrel, knots on the barrel wall, foreign, misaligned, missing, or extra piston, dirt on the piston, deformed ribs on the piston, pistons in impermissible positions, flow marks on the barrel wall, voids in the flange, barrel, or luer lock plastic, imperfectly molded cartridges, missing, cut, misaligned, loose, or damaged luer lock caps, etc. Examples of defects related to the interaction between the cartridge and cartridge contents may include liquid leaking through the piston, liquid in the piston ribs, etc. Figure 3B, described below, illustrates various components of an exemplary cartridge.
[0027] Non-limiting examples of defects related to the vial may include cracks, chips, scratches, and / or scrapes in the body, air lines in the glass of the body, discontinuities in the glass of the body, stains on the inside or outside (or internal) of the body, stuck glass on the body, knots in the body wall, flow marks in the body wall, missing, misaligned, loose, protruding, or damaged crimps, missing, misaligned, loose, or damaged flip caps, etc. Examples of defects related to the interaction between the vial and the vial contents may include leakage of liquid through the crimp or cap, etc. Figure 3C, described below, illustrates various components of an exemplary vial.
[0028] Non-limiting examples of defects associated with container contents (e.g., the contents of a syringe, cartridge, vial, or other container type) include foreign objects suspended in the liquid contents, foreign objects on the plunger dome, piston dome, or vial floor, discolored liquid or cake, cracks, dispersions, or other abnormally dispersed / formed cake, cloudy liquid, high or low fill levels, etc. "Foreign" particles may include, for example, fibers, rubber, metal, stones, plastic fragments, hair, etc. In some embodiments, air bubbles are considered harmless and are not considered to be defects.
[0029] Non-limiting examples of other types of features that may be shown in images of feature image library 142 may include menisci of different shapes and / or different locations, plungers of different types and / or different locations, air bubbles of different sizes and / or shapes and / or different locations within a container, different void sizes within a container, different sizes, shapes, and / or locations of irregularities in glass or another translucent material, etc.
[0030] During operation, computer system 104 stores sample images collected by visual inspection system 102 (possibly after cropping and / or other image preprocessing by computer system 104), as well as synthetic images generated by library expansion module 124, and possibly real and / or synthetic images from one or more other sources, in training image library 140. AVI neural network module 120 then trains an AVI neural network using at least some of the sample images in training image library 140 and validates the trained AVI neural network using other images in library 140 (or in another library not shown in FIG. 1 ). As used herein, the terms “training,” “validating,” or “qualifying” a neural network encompass both directly executing software that executes the neural network (e.g., by instructing or requesting a remote server to train the neural network or execute a trained neural network) and initiating execution of the neural network. In some embodiments, for example, computer system 104 may "train" the neural network by accessing a remote server that includes AVI neural network module 120 (e.g., by accessing a web service supported by the remote server).
[0031] The operation of each of the modules 120-126 is discussed in further detail below with reference to various other figure elements.
[0032] FIG. 2 illustrates an exemplary visual inspection system 200 that may be used as the visual inspection system 102 of FIG. 1 in a pharmaceutical application. The visual inspection system 200 includes a camera 202, a lens 204, forward-tilting light sources 206 a and 206 b, rear-tilting light sources 208 a and 208 b, a backlight source 210, and an agitation mechanism 212. The camera 202 captures one or more images of a container 214 (e.g., a syringe, vial, cartridge, or other suitable type of container), which is held by the agitation mechanism 212 and illuminated by light sources 206, 208, and / or 210 (e.g., the VIS control module 122 activates different light sources sequentially or simultaneously for different images). The visual inspection system 200 may include additional or fewer light sources (e.g., by omitting the backlight source 210). The container 214 may hold, for example, a liquid or lyophilized pharmaceutical product.
[0033] Camera 202 may be a high-performance industrial camera or a smart camera, and lens 204 may be, for example, a high-fidelity telecentric lens. In one embodiment, camera 202 includes a charge-coupled device (CCD) sensor. For example, camera 202 may be a Basler® pilot piA2400-17gm monochrome area-scan CCD industrial camera with a resolution of 2448 x 2050 pixels. As used herein, the term "camera" may refer to any suitable type of imaging device (e.g., a camera that captures the human visible portion of the frequency spectrum, an infrared camera, etc.).
[0034] Different light sources 206, 208, and 210 may be used to collect images for detecting different categories of defects. For example, forward-tilted light sources 206a and 206b may be used to detect reflective particles or other reflective defects, rear-tilted light sources 208a and 208b may be used for particles generally, and backlight light source 210 may be used to detect opaque particles and / or to detect incorrect dimensions and / or other defects in a container (e.g., container 214). Light sources 206 and 208 may include CCS® LDL2-74X30RD bar LEDs, and backlight light source 210 may be, for example, a CCS® TH-83X75RD backlight.
[0035] Stirring mechanism 212 may include a chuck or other means for holding and rotating (e.g., pivoting) a container, such as container 214. For example, stirring mechanism 212 may include an Animatics® SM23165D SmartMotor with a spring-loaded chuck that securely attaches each container (e.g., a syringe) to the motor.
[0036] While the visual inspection system 200 may be suitable for generating container images to train and / or validate one or more AVI neural networks, the ability to detect defects across a wide range of categories may require different perspectives. Thus, in some implementations, the visual inspection system 102 of FIG. 1 may instead be a multi-camera system. In yet other implementations, the visual inspection system 102 of FIG. 1 may include a line-scan camera and rotate the sample (e.g., container) to capture each image. Furthermore, automated handling / transportation of the sample may be desirable to quickly obtain a much larger set of training images. The visual inspection system 102 may be, for example, any of the visual inspections shown and / or described in U.S. Provisional Patent Application No. 63 / 020,232 (filed May 5, 2020, entitled “Deep Learning Platforms for Automated Visual Inspection”), the entirety of which is incorporated herein by reference, or any other visual inspection system suitable for any type of product. For example, in an automotive context, the visual inspection system 200 may include a conveyor belt with illumination sources and multiple cameras mounted above and / or around specific conveyor belt stations.
[0037] 3A-3C illustrate various exemplary container types that may be used as samples to be imaged by the visual inspection system 102 of FIG. 1 or the visual inspection system 200 of FIG. 2 in a particular pharmaceutical context. Referring first to FIG. 3A, an exemplary syringe 300 includes a hollow barrel 302, a flange 304, a plunger 306 that provides a movable fluid seal within the interior of the barrel 302, and a needle shield 308 that covers the syringe needle (not shown in FIG. 3A). For example, the barrel 302 and flange 304 may be formed from glass and / or plastic, and the plunger 306 may be formed from rubber and / or plastic. The needle shield 308 is separated by a gap 312 by a shoulder 310 of the syringe 300. The syringe 300 contains a liquid (e.g., a pharmaceutical product) 314 within the barrel 302 and above the plunger 306. The top of the liquid 314 forms a meniscus 316, above which is a void 318.
[0038] 3B, an exemplary cartridge 320 includes a hollow barrel 322, a flange 324, a piston 326 that provides a movable fluid seal within the barrel 322, and a luer lock 328. For example, the barrel 322, flange 324, and / or luer lock 328 may be formed from glass and / or plastic, and the piston 326 may be formed from rubber and / or plastic. The cartridge 320 includes a liquid (e.g., a pharmaceutical product) 330 within the barrel 322 and above the piston 326. The top of the liquid 330 forms a meniscus 332, with a void 334 above it.
[0039] 3C, an exemplary vial 340 includes a hollow body 342 and a neck 344, with the transition between the two forming a shoulder 346. At the bottom of the vial 340, the body 342 transitions into a heel 348. A crimp 350 includes a stopper (not visible in FIG. 3C) that provides a fluid seal at the top of the vial 340, and a flip cap 352 covers the crimp 350. For example, the body 342, neck 344, shoulder 346, and heel 348 may be formed from glass and / or plastic, the crimp 350 may be formed from metal, and the flip cap 352 may be formed from plastic. The vial 340 may contain a liquid (e.g., a pharmaceutical product) 354 within the body 342. The top of the liquid 354 may form a meniscus 356 (e.g., a very slightly curved meniscus if the body 342 has a relatively large diameter), above which there is a void 358. In other embodiments, the liquid 354 is instead a solid material within the vial 340. For example, the vial 340 may contain a lyophilized (freeze-dried) pharmaceutical product 354, also referred to as a "cake."
[0040] For example, various image augmentation techniques that may be implemented by library expansion module 124 (executed by processing unit 110) are now described. Referring initially to FIG. 4A , module 124 may implement an arithmetic transposition algorithm 400 to add features (e.g., defects) to an original (e.g., real) image with pixel-level realism. While FIG. 4A describes algorithm 400 with reference to a “container” image, and in particular a glass container, it will be understood that module 124 may instead use algorithm 400 to augment images of other types of samples (e.g., plastic containers, vehicle body parts, etc.).
[0041] Initially, in block 402, module 124 loads into memory (e.g., memory unit 114) a defect image and an image of a container without defects shown in the defect image. The container image (e.g., a syringe, cartridge, or vial similar to one of the containers shown in FIGS. 3A-3C) may be a real image captured, for example, by visual inspection system 102 of FIG. 1 or visual inspection system 200 of FIG. 2. Depending on the implementation, the real image may have been processed in other ways (e.g., cropped, filtered, etc.) before block 402. The defect image may be, for example, a specific type of defect (e.g., a scratch, a crack, a stain, a foreign object, a deformed plunger, a broken cake, etc.) that module 124 retrieved from feature image library 142.
[0042] In block 404, module 124 converts the defect image and the container image into respective two-dimensional numerical matrices, referred to herein as a “defect matrix” and a “container image matrix,” respectively. Each of these numerical matrices may include one matrix element for each pixel in the corresponding image, with each matrix element having a numerical value representing the (grayscale) intensity value of the corresponding pixel. For example, for a typical industrial camera having an 8-bit format, each matrix element may represent an intensity value from 0 (black) to 255 (white). For example, in an implementation where the container is backlit, areas of the container image showing only glass and clear fluid may have relatively high intensity values, while areas of the container image showing defects may have relatively low intensity values. However, algorithm 400 may be useful in other scenarios as long as the intensity levels of the indicated defects are sufficiently different from the intensity levels of glass / fluid areas shown without defects. Other numerical values may be used for other grayscale resolutions, or the matrices may have more dimensions (e.g., if the camera generates red-green-blue (RGB) color values). 5 illustrates an exemplary operation by module 124 to convert a feature (crack) image 500 having grayscale pixels 502 into a feature matrix 504. For clarity, FIG. 5 illustrates only a portion of the pixels 502 in feature image 500 and only a portion of the corresponding feature matrix 504.
[0043] The two-dimensional matrix generated in block 404 for a vessel image of pixel size m×n can be expressed as the following m×n matrix:
number
number
[0044] In block 406, the library expansion module 124 sets restrictions on where defects can be placed in the container image. For example, the module 124 may disallow placement of defects in areas of the container that have large discontinuities in intensity and / or appearance, such as by disallowing placement in areas outside of translucent fluid in a transparent container. In other embodiments, defects can be placed anywhere on the sample.
[0045] In block 408, module 124 identifies a "surrogate" region within the container image, within any limits set in block 406. The surrogate region is the region into which defects will be transferred and is therefore the same size as the defect image. Module 124 may identify the surrogate region using a random process (e.g., randomly selecting x- and y-coordinates within the limits set in block 406) or may set the surrogate region at a predetermined location (e.g., in an implementation in which module 124 steps through different displaced positions at regular or irregular intervals / spaces over multiple iterations of algorithm 400).
[0046] In block 410, module 124 generates a surrogate region matrix corresponding to the surrogate region of the container image. The matrix may be formed by converting pixel intensities in the surrogate region of the original container image to numerical values, or may simply be formed by directly copying numerical values from the corresponding portion of the container image matrix generated in block 404. In either case, the surrogate region matrix corresponds to the exact location / region of the container image where the defect will be transposed, and is identical in size and shape (i.e., number of rows and columns) to the defect matrix. Thus, the surrogate region matrix may have the following form:
number
[0047] In block 412, for each row of the defect matrix, module 124 generates a histogram of element values. An exemplary defect histogram 450 for a single row of the defect matrix is shown in FIG. 4B. In histogram 450, a first peak 452 corresponds to relatively low-intensity pixel values for regions of the defect image that show the defect itself, a second peak 454 corresponds to relatively medium-intensity pixel values for regions of the defect image that show only glass / fluid (no defects), and a third peak 456 corresponds to relatively high-intensity pixel values for regions of the defect image that show light reflections from the defect. Careful selection of the defect image size is important to ensure that histogram 450 includes peaks 454. In particular, the defect image loaded in block 402 should be large enough to capture at least some glass regions (i.e., no defects) across all rows of the defect image.
[0048] For each row of the defect matrix, module 124 also (at block 412) identifies a peak portion (e.g., peak portion 454 in histogram 450) corresponding to the indicated glass without defects and normalizes the element values of that row of the defect matrix to the center of that peak portion. In some implementations, the dimensions of the defect image are selected so that the peak portion with the largest peak corresponds to the glass / defect-free region of the defect image. In these implementations, module 124 can identify the peak portion corresponding to the indicated glass (without defects) by selecting the peak portion with the largest peak value. Module 124 can determine the “center” of the peak portion in various ways, depending on the implementation. For example, module 124 can determine the low-side and high-side intensity values of the peak portion (shown in exemplary histogram 450 as low-side value (LSV) 457 and high-side value (HSV) 458, respectively) and then calculate the average of the two (i.e., center=(HSV−LSV) / 2). Alternatively, module 124 can calculate the center as the median intensity value, or the intensity value corresponding to the peak of the peak portion, etc. The HSV and LSV values of the defect image may be fairly close, for example, about 8-10 grayscale levels apart.
[0049] To normalize the defect matrix, module 124 subtracts a center value from each element value in a row. An example of this is shown in FIG. 4C , where a defect image having histogram 450 has been normalized so that the normalized defect matrix has histogram 460. As can be seen in FIG. 4C , in this example, peak portion 452 has been transformed into peak portion 462, which contains only negative values; peak portion 454 has been transformed into peak portion 464, which is centered around element value 0; and peak portion 456 has been transformed into peak portion 466, which contains only positive values. It is understood that module 124 does not necessarily generate histogram 460 when performing algorithm 400. In effect, the normalized defect matrix is a “flattened” version of the defect matrix, where values are canceled out by the surrounding glass (and possibly fluid, etc.) while retaining information representative of the defects themselves. When performed for all rows, the normalized defect matrix may be expressed as follows:
number
[0050] In block 414, module 124 generates a similar histogram for each row of the surrogate region matrix, identifies a peak corresponding to the glass / fluid represented in the surrogate region, and records the low-side and high-side values of the peak. In implementations / scenarios where the container image does not show any defects, there may be only one peak in the histogram (e.g., similar to peak 450 in LSV 457 and HSV 458). Because the lighting (and possibly other) conditions when the defect image and the container image are captured are not exactly the same, the peak identified in block 414 will differ in at least some respects from the defect image peak identified in block 412.
[0051] It is understood that algorithm 400 may be performed row-wise, as described above, or column-wise. Performing the operations of blocks 412 and 414 row-wise or column-wise may be particularly advantageous when a cylindrical container is positioned orthogonal to the camera, with the container's central / long axis extending horizontally or vertically across the container image. In such a configuration, depending on the type and location of illumination, changes in appearance tend to be more abrupt in one direction (across the diameter or width of the container) and less abrupt in another direction (along the container's long axis); therefore, less information is lost by performing normalization, etc., on each row or column (i.e., whichever corresponds to the direction with less change). In some implementations (e.g., when imaging a vial from the bottom side), blocks 412 and 414 may involve other operations, such as averaging values within a two-dimensional region (e.g., 2×2, 4×4, etc.) of the surrogate region matrix.
[0052] In blocks 416-420, module 124 iteratively performs a comparison for each element of the defect matrix (e.g., element D 11 414A, the normalized defect matrix is mapped onto the surrogate region of the container image matrix (by scanning the defect matrix starting from ). For a given element of the normalized defect matrix, in block 416, module 124 adds the value of that element to the value of the corresponding element of the surrogate region matrix and determines whether the resulting sum falls between the low-side and high-side values of the corresponding row (as determined in block 414). If so, then in block 418A, module 124 retains the original value of the corresponding element in the surrogate region of the container image matrix.
[0053] If not, then in block 418B, module 124 adds the normalized defect matrix element value to the value of the corresponding element of the vessel image matrix. For example, element N 11 If is outside the range [LSV,HSV], the module 124 converts the corresponding element in the vessel image to (N 11 +S 11) for each remaining element of the normalized defect matrix. As indicated in block 420, module 124 repeats block 416 (and block 418A or block 418B, as appropriate) for each remaining element of the normalized defect matrix. In block 422, module 124 verifies that all values in the modified container image (at least the surrogate region) are valid bitmap values (e.g., between 0 and 255 if an 8-bit format is used), and in block 424, module 124 converts the modified container image matrix to a bitmap image and stores the resulting “defective” container image (e.g., in training image library 140). The net effect of blocks 416-420 is to “catch” or retain defect image pixels that are less intense (darker) than the glass (or other translucent material) level in the container image, as well as pixels that are more intense (brighter / whiter) than the glass level (e.g., due to reflections in the defect).
[0054] It is understood that the various blocks described above for algorithm 400 may differ in other implementations, including in ways other than (or in addition to) the various options described above. By way of example only, the loop of blocks 416-420 may involve first merging the normalized defect matrix with the surrogate area matrix (element-by-element, as described above for the container image matrix) to form a permutation matrix, and then replacing corresponding regions of the container image matrix with the permutation matrix (i.e., rather than directly modifying the entire container image matrix). As another example, blocks 416A and 416B may instead operate to modify the normalized defect matrix (i.e., by changing element values to zero each time block 418A is executed), after which the modified version of the normalized defect matrix is added to the surrogate region of the container image matrix. Furthermore, algorithm 400 may omit one or more operations described above (e.g., block 406) and / or include additional operations not described above.
[0055] In some embodiments and / or scenarios, algorithm 400 includes rotating and / or scaling / resizing the defect image (loaded in block 402) or a numerical matrix derived from the defect image (in block 404) before transposing the defect onto the surrogate region of the container image. For example, the rotation and / or resizing of the defect image or numerical matrix may occur any time before block 412 (e.g., immediately before any one of blocks 410, 408, 406, and 404). Rotation may be performed, for example, with respect to a center point or central pixel of the defect image or numerical matrix. Resizing may include expanding or contracting the defect image or numerical matrix along one or two axes (e.g., along the axes of the defect image or along the major and minor axes of the depicted defect, etc.). In general, scaling / resizing an image involves mapping a group of pixels onto a single pixel (shrinking) or mapping a single pixel onto a group of pixels (expanding / stretching). It is understood that when performing operations on a numerical matrix derived from a defect image, similar operations are required on matrix elements rather than pixels. Once the defect image or numerical matrix is rotated and / or resized, the remainder of algorithm 400 may remain unchanged (i.e., occur in the same manner as described above and may be agnostic as to whether rotation and / or resizing occurred).
[0056] Rotation and / or resizing (e.g., by library expansion module 124 implementing arithmetic transposition algorithm 400) can help increase the size and diversity of feature image library 142 far beyond what would otherwise be possible with a fixed set of defect images. Rotation can be particularly useful in use cases where (1) the imaged container has significant rotational symmetry (e.g., the container has a circular or semicircular surface that is imaged during inspection) and (2) the imaged defect is of a type that tends to have visual characteristics that depend on that symmetry. For example, in a circular or near-circular bottom of a glass vial, some cracks may tend to propagate generally in a direction from the center of the circle to the periphery of the circle, or vice versa. For example, library expansion module 124 can rotate a crack or other defect so that the axes of the defect image are aligned with the rotational position of the surrogate region into which the defect is being transposed. More specifically, the amount of rotation may depend on both the rotation of the defect in the original defect image and the desired rotation (e.g., the rotation corresponding to the surrogate region into which the defect is being transposed).
[0057] Any suitable technique, such as nearest neighbor, bilinear, high-quality bilinear, bicubic, or high-quality bicubic, may be used to achieve the pixel (or matrix element) mapping required for the desired rotation and / or resizing. Of the five exemplary techniques listed above, nearest neighbor is a lower-quality technique, and high-quality bicubic is the highest-quality technique. However, given that the goal is to achieve an image quality of the rotated and / or resized defects that closely matches the image quality provided by the imaging system (e.g., visual inspection system 102) that will be used for inspection, the highest-quality technique may not be optimal. Manual user review may be performed to compare the output of different techniques, such as the five listed above, and select the technique that is best in a qualitative / subjective sense. In some implementations, high-quality bicubic is used or is used as the default setting.
[0058] Algorithm 400 (with or without optional rotation and / or resizing) can be repeated for any number of different "good" images and any number of "defective" images in any desired combination (e.g., applying each of L defect images to each of M good container images at each of N locations to generate an L x M x N composite image based on the M good container images in training image library 140). Thus, for example, 10 defect images, 1,000 good container images, and 10 defect locations per defect type may result in 100,000 defect images. The locations / positions at which defects are displaced for any particular good container image may be predetermined or may be randomly determined (e.g., by module 124).
[0059] Algorithm 400 can also perform very well in situations where a defect is transposed onto a surrogate region that includes sharp contrasts or transitions in pixel intensity levels due to one or more features. For example, algorithm 400 can perform well even when the surrogate region of a glass syringe includes the meniscus and regions on either side of the meniscus (i.e., air and fluid, respectively). Algorithm 400 can also handle certain other situations where the surrogate region is very different from the region surrounding the defect in the defect image. For example, algorithm 400 can perform well when transposing a defect from a defect image of a glass syringe filled with a transparent fluid onto an image of a vial in a surrogate region where the vial is filled with an opaque lyophilized cake. However, for some use cases or scenarios, it may be beneficial to modify algorithm 400. If the surrogate region of the container image shows a transition between two very different regions (e.g., between the glass / air portion of a vial image and the lyophilized cake portion), for example, module 124 can divide the surrogate region matrix into multiple portions (e.g., two matrices of the same size or different sizes), or simply form two or more surrogate region matrices in the first example. Corresponding portions of the defect image may then be individually transposed onto the different surrogate regions using different instances of algorithm 400 as discussed above.
[0060] In some implementations, defects and / or other features depicted in images of feature image library 142 can be morphed in one or more ways before module 124 uses algorithm 400 to add those features to the original image. In this way, module 124 can effectively increase the size and variability of feature image library 142, and thus the size and variability of training image library 140. For example, module 124 can morph defects and / or other features by applying rotation, scaling / stretching (in one or two dimensions), skewing, and / or other transformations. Additionally or alternatively, depicted features may be modified in more complex and / or subtle ways. For example, module 124 can fit defects (e.g., cracks) into different arcs or more complex crack structures (e.g., into each of several different branching patterns). By its nature, pixel-based algorithm 400 is well-equipped to handle this type of fine feature control / modification.
[0061] The composite images generated using the arithmetic transposition algorithm 400 of FIG. 5 can be highly realistic, as can be seen in FIG. 6. FIG. 6 compares a real image 600 of a syringe with a real-world crack that was manually generated with a synthetic image 602 of a syringe with a crack that was artificially generated using the algorithm 400. Furthermore, the “realism” of the composite images can be extended down to the pixel level. FIG. 7 provides a pixel-level comparison corresponding to images 600 and 602 of FIG. 6. Specifically, image portion 700A is a close-up of a real-world defect in container image 600, and image portion 702A is a close-up of an artificial defect in container image 602. Image portion 700B is a further close-up of image portion 700A, and image portion 702B is a further close-up of image portion 702A. As can be seen in image portions 700B and 702B, there are no readily observable pixel-level artifacts or other dissimilarities created by transposing the defect.
[0062] Without this pixel-level realism, the AVI neural network may focus on "wrong" characteristics (e.g., pixel-level artifacts) when determining that a composite image is defective. While container materials (e.g., glass or plastic) appear to the naked eye as uniform surfaces, in reality, lighting and container material properties (e.g., container curvature) result in pixel-to-pixel variations, such that each surrogate region on a given container image differs in at least some respects from all other possible surrogate regions. Furthermore, the conditions / materials (e.g., lighting and container material / shape) used to capture a defective image may exhibit even greater variation compared to the conditions / materials used to capture a "good" container image. A potential example of this is shown in FIG. 8 , which shows a composite composite image 800 with both a first transposed defect 802 and a second transposed defect 804. The first transposed defect 802 was created using a simple conventional technique of directly superimposing the defect image onto the original container image, while the second transposed defect 804 was created using the arithmetic transposition algorithm 400. 8, the boundary of the defect image corresponding to the first displaced defect 802 is clearly visible. An AVI neural network trained using a composite image having a defect such as the first displaced defect 802 may simply look for similar boundaries when inspecting a container, for example, which may result in a large number of false negatives and / or other inaccuracies.
[0063] 9A shows various other composite images with added defects, labeled 900-910, generated using an implementation of arithmetic transposition algorithm 400. In each case, the portion of the syringe image showing the defect blends seamlessly with the surrounding portions of the image, whether the image is viewed at a macroscopic level or at a pixel level.
[0064] FIG. 9B shows a collection of example crack and defect images 920, any of which may be used as input to the arithmetic transposition algorithm 400. In some implementations, as noted above, the arithmetic transposition algorithm 400 may include rotating and / or resizing a given defect image (or corresponding numerical matrix) before executing the remainder of the algorithm 400. When rotation is desired, it is generally important to know the rotation corresponding to the original defect image. For example, for the example crack and defect images 920, the rotation / angle corresponding to the original image is included in the filename itself (shown immediately below each image in FIG. 9B). Thus, for example, “250_crack0002” could be a particular crack at a 250-degree rotation (such that a 70-degree counterclockwise rotation is required to position the crack where a 180-degree rotation is desired), “270_crack0003” could be another crack at a 270-degree rotation (such that a 90-degree counterclockwise rotation is required to position the crack where a 180-degree rotation is desired), and so on. The library expansion module 124 can calculate the degree of rotation to apply based on this indicated original rotation and the desired rotation (e.g., the rotation corresponding to the angular position of the surrogate region to which the defect is displaced).
[0065] The arithmetic transposition algorithm 400 can be implemented in most high-level languages, such as C++, the .NET environment, etc. Depending on the processing power of the processing unit 110, the algorithm 400 can potentially generate thousands of synthetic images in under 15 minutes, although rotation and / or resizing typically increase these times. However, because training images do not need to be generated in real time for most applications, execution time (even with rotation and / or resizing) is generally not a significant issue.
[0066] As described in U.S. Provisional Patent Application No. 63 / 020,232, various image processing techniques can be used to measure key metrics for each available image, allowing for careful curation of a training image library, such as training image library 140. During the development of the above-described arithmetic transposition algorithm 400, it was discovered that careful control of certain parameters can be important. For example, when considering a 1 ml glass syringe, the position of the liquid meniscus and plunger (e.g., rubber plunger) in the image can be important attributes that can vary from image to image. If synthetic images are created with all the same “good” vessel images (or with a set of good vessel images that are too small and / or too similar), subsequent training of the deep learning AVI model may be compromised by bias resulting from a lack of image variability.
[0067] By using key image metrics, a library of "good" images to be expanded (e.g., using algorithm 400) can be carefully selected so that these biases are reduced or avoided. Such metrics can also be used to blend the training image library, so that the resulting synthetic library not only contains an appropriate balance of real and synthetic images, but also naturally displays the variance of each of the key metrics.
[0068] Various experiments were conducted to evaluate the quality of synthetic images generated using algorithm 400, including the robustness of AVI deep learning models trained on such images. Four datasets, each with approximately 300 images, were used for these experiments: (1) a “real defect-free” image set, which were real images of syringes without visible defects captured by a Cartesian robot-based system in a laboratory setting; (2) a “real defect” image set, which were real images of syringes with cracks of different sizes in different locations, also captured by a Cartesian robot-based system in a laboratory setting; (3) a “synthetic defect-free” image set, which were synthetic images created by removing the indicated cracks from the real defect images without changing the positions of the plunger and meniscus; and (4) a “synthetic defect” image set, which were synthetic images created by adding depictions of cracks to the real defect-free images and randomly positioning them in the X and Y directions. The synthetic defect images were generated using an implementation of algorithm 400. The syringes in the real defect-free and real defect-containing images had meniscuses in different locations.
[0069] The AVI deep learning model was trained using different combinations of the proportion of images from the real dataset and the augmented dataset (0%, 50%, or 100%). For each combination, two image libraries were blended: a good (defect-free) image library and a defective image library, each with approximately 300 images. During training, each of these two libraries was divided into three parts: 70% of the images were used for training, 20% for validation, and 10% for the test dataset. A pre-trained ResNet50 algorithm was used to train the model using HALCON® software to classify input images into defective or non-defective classes. After training the deep learning model, its performance was evaluated using the test dataset. When the model was trained with 0% real images (i.e., 100% synthetic images), the accuracy of the augmented test set was observed to be higher than that of the real dataset. When the model was trained with 100% real images (i.e., 0% synthetic images), the accuracy of the real dataset was higher than that of the augmented dataset. The model was trained using 50% real images and 50% synthetic images, and the accuracy was similar and high for both the real and augmented datasets. From these experiments, it was concluded that increasing the proportion of either real or synthetic images in the training dataset correspondingly increases the accuracy of the deep learning model for the respective (real or augmented) dataset.
[0070] When the model was trained on 100% real images, one possible reason for the lower accuracy of the model for the synthetic / augmented test images could be that the meniscus of the syringe in the training and test image sets was different. A model trained on 0% real images and tested only on real images could misclassify test images due to the different meniscus. Similarly, when trained on 100% real images and tested using only synthetic images, the model could misclassify test images due to the different meniscus. These misclassified images were evaluated by visualizing heat maps generated using the Gradient Class Activation Map (Grad-CAM) algorithm. This type of heat map is discussed in more detail in U.S. Provisional Patent Application No. 63 / 020,232. In such cases, the image augmentation techniques discussed herein can be used to improve classifier performance by adding variability to the meniscus in the training images.
[0071] After the model was trained and the above tests indicated that it was properly trained, a "final test" phase was performed. For this phase, the same four general types of datasets ("No Real Defects," "Real Defects," "No Synthetic Defects," and "Synthetic Defects") were again used, but all images were obtained from a different source (i.e., all images were of a different product than those used in the training / validation / test phase), and all images were used only to test model performance (i.e., none of the images were used to train the model). A similar trend was observed for this second phase: a higher proportion of real images increased model accuracy for the real "final test" images, and a higher proportion of synthetic images increased model accuracy for the synthetic "final test" images.
[0072] FIG. 10 shows heatmaps 1000, 1002, and 1004 generated by various Grad-CAMs used to evaluate the effectiveness of the synthetic images. Heatmap 1000 reflects a “true positive,” i.e., a case in which the AVI neural network correctly identified a digitally added crack. As seen in FIG. 10, the pixels associated with the crack were the pixels the AVI neural network relied on most when making its “defect” inference. However, heatmap 1002 reflects a “false positive,” in which the AVI neural network classified the synthetic image as defective but for the wrong reason (i.e., by focusing on an area away from the digitally added crack). Heatmap 1004 reflects a “false negative,” in which the AVI neural network failed to classify the synthetic image as defective because the model over-focused on the meniscus area. This misclassification is a result of the synthesized “defective” training image having a meniscus similar to the “defect-free” test image. This is most likely to occur when training on 100% real images and then running the model on synthetic images, or when training on 100% synthetic images and then running the model on real images. Instead, training on something like 50% real and 50% synthetic images dramatically reduces such artifacts.
[0073] The AVI neural network performance was also measured by generating a confusion matrix for the AVI model when using different combinations of real and synthetic images as training data. When the AVI model was trained on 100% synthetic images, the model performance for the 100% synthetic image set was as follows:
[0074] [Table 1]
[0075] When the AVI model was trained on 50% real and 50% synthetic images, the model performance on the 100% synthetic image set was as follows:
[0076] [Table 2]
[0077] When the AVI model was trained on 100% real images, the model performance on the 100% synthetic image set was as follows:
[0078] [Table 3]
[0079] When the AVI model was trained on 100% synthetic images, the model performance on the 100% real image set was as follows:
[0080] [Table 4]
[0081] When the AVI model was trained on 50% real and 50% synthetic images, the model performance on the 100% real image set was as follows:
[0082] [Table 5]
[0083] When the AVI model was trained on 100% real images, the model performance on the 100% real image set was as follows:
[0084] [Table 6]
[0085] These results are also reflected in FIG. 11 , which is a plot 1100 showing AVI neural network performance for different combinations of synthetic and real images in the training and test image sets. In plot 1100, the x-axis represents the percentage of real images in the training set, with the remainder being synthetic / augmented images, and the y-axis represents the percentage accuracy of the trained AVI model. Trace 1102 corresponds to tests performed on 100% real images, and trace 1104 corresponds to tests performed on 100% synthetic images. As can be seen from plot 1100 and the confusion matrix above, a mixture of 50% real images and approximately 50% synthetic images (e.g., in training image library 140) appears to be optimal (approximately 98% accuracy). Of course, the density of data points in plot 1100 may mean that the optimum point is somewhat above or below the 50% real image point. For example, if a 5-10% lower proportion of real training images results in something very close to 98% accuracy, it may be desirable to accept a small degradation in performance (when tested on real images) in order to save the cost / time of developing a training image library with a higher proportion of synthetic images.
[0086] The above discussion has primarily focused on generating synthetic “flawed” images, i.e., augmenting “good” real images by adding artificial but realistically presented defects. However, in some cases, it may be advantageous to create synthetic “good” images from real images that exhibit defects or anomalies. This can further expand the training image library and also help balance the characteristics of the “flawed” and “flaw-free” images in the training image library. In particular, removing defects can reduce non-causal associations by the AVI model by providing complementary counters to images that exhibit defects. This encourages the AVI model to focus on the appropriate regions of interest to identify causal associations that may be very subtle in some cases.
[0087] In some implementations, defect (or other feature) removal is performed on a subset of images that exhibit the defect of interest, after which both the synthetic (defect-free) image and the corresponding original (defective) image are included in a training set (e.g., in training image library 140). AVI classification models trained on good images, independent of the defective samples, where approximately 10% of the training images are synthetic "good" images created from defective images, have been shown to match or exceed the causal prediction performance of AVI models trained on good images sourced entirely from defective samples where no defect artifacts are visible in the images.
[0088] More general feature removal (as opposed to only defect removal) can be used to provide more focused classification. For example, if the original images in the training set show specific regions or areas of interest (e.g., the meniscus, which can vary in appearance and position), such regions can be replaced (e.g., by removing or modifying the identified characteristics of those regions) and the edited images can be added as complementary training images. This may be preferable to cropping (e.g., cropping out a portion of the syringe image showing the meniscus), for example, if the AVI model requires a specific input size and / or if there are multiple scattered regions of interest.
[0089] To remove the indicated defect or other feature from the original image, different digital “inpainting” techniques are described herein. In some implementations, module 124 removes the image feature by first masking the defect or other feature (e.g., uniformly setting all pixels corresponding to the feature region to a minimum or maximum intensity) and then iteratively searching the masked image for a region that best “fits” the hole (masked portion) by matching surrounding pixel statistics. More specifically, module 124 can determine correspondences between (1) portions of the image adjacent to the masked region (e.g., patches) and (2) other portions of the image outside the masked region. For example, module 124 can inpaint the masked region using a patch-matching algorithm. If the unmasked region of the image does not exhibit the same feature (e.g., the same defect) as the masked region, module 124 will remove that feature when filling the masked region.
[0090] This inpainting technique can generally produce "smooth" and realistic-looking results. However, this technique is limited by available image statistics and has no concept of image themes or semantics. Therefore, some synthesized images may not be subtly or holistically representative of actual "good" images. To address these concerns, some implementations use deep learning-based inpainting. These techniques use neural networks to map complex relationships between input images and output labels. Such models can learn higher-level image themes and identify meaningful correlations that provide continuity in the augmented image.
[0091] In some deep learning implementations, module 124 inpaints the image using a partial convolutional model. A partial convolutional model performs convolutions across the entire image, which adds pixel noise and variability to the composite (inpainted) image, thus slightly distinguishing the composite image from the original, even beyond the inpainted region. Using a composite image with this pixel noise / variability (e.g., by AVI neural network module 120) to train the AVI model can help prevent overfitting of the model, as the additional variability prevents the model from depicting correlations inherent in the overlay. This allows the AVI model to "understand" the total image population better, rather than just a specific subset of that population. The result is a more efficiently trained and focused AVI deep learning model.
[0092] FIG. 12 illustrates an exemplary partially convolutional model 1200 that module 124 can use to generate a composite image. The general structure of model 1200, known as a "U-Net" architecture, has been used in image segmentation applications. In model 1200, an input image and mask pair 1202 are input (as two separate inputs with the same dimensions) to encoder 1204 of model 1200. In the example shown in FIG. 12, the image and mask of input pair 1202 both have 512×512 pixels / elements and both have three dimensions per pixel / element (to represent red, green, and blue (RGB) values). In other implementations, the image and mask of input pair 1202 may be larger or smaller in width and height (e.g., 256×256), and may have more or fewer than three pixel dimensions (e.g., one-dimensional if a grayscale image is used).
[0093] During training, when module 124 inputs a particular input and mask as input pair 1202, model 1200 dots the image with the mask (i.e., applies the mask to the image) to form training samples, while the original image (i.e., the image in input pair 1202) serves as the target image. In the first stage of encoder 1204, model 1200 applies the masked version of the input image and the mask itself as separate inputs to a two-dimensional convolutional layer to generate an image output and a mask output, respectively. The mask output at each stage may be clipped to the range [0, 1]. Model 1200 dots the image output and mask output and feeds the dotted image output and mask output as separate inputs to the next two-dimensional convolutional layer. Model 1200 iteratively repeats this process until no convolutional layers remain in encoder 1204. In each successive convolutional layer, the pixel / element dimension may increase up to a certain value (512 in the example of FIG. 12), while the size of the masked image and mask decreases until they reach a sufficiently small size (2×2 in the example of FIG. 12). The encoder 1204 has N 2D convolutional layers, where N is any suitable integer greater than 1 and is an adjustable hyperparameter. Other adjustable hyperparameters of the model 1200 may include kernel size, stride, and padding.
[0094] After model 1200 passes the (masked) image and mask through encoder 1204, model 1200 passes the (now smaller, but higher-dimensional) masked image and mask through a transposed convolutional layer in decoder 1206. Decoder 1206 contains the same number (N) of layers as encoder 1204 and restores the image and mask to their original size / dimensions. Before each transposed layer in decoder 1206, model 1200 concatenates the image and mask from the previous layer (i.e., from the last convolutional layer of encoder 1204 or from a previous transposed layer of decoder 1206) with the output of the corresponding convolutional layer in encoder 1204, as shown in FIG. 12.
[0095] The decoder 1206 outputs an output pair 1208, which includes a reconstructed (output) image and a corresponding mask. For training, as described above, the original image serves as a target image to which the module 124 compares the image in the output pair 1208 at each iteration. The module 124 can train the model 1200 by attempting to minimize the following six losses: Effective loss: pixel loss in the area outside the mask. The module 124 can calculate this loss by summing up the pixel value differences between the input / original image and the output / reconstructed image. - Hole loss: pixel loss in masked areas. Perceptual loss: A higher-level feature loss that module 124 can compute using a separately trained (pre-trained) VGG16 model. The VGG16 model may be pre-trained to classify samples with and without relevant features (e.g., defects). During training of model 1200, module 124 can compute the perceptual loss by feeding the original and reconstructed images to the pre-trained VGG16 model and taking the difference of three max pooling layers in the VGG16 model for the original and reconstructed images. Style Loss 1: Module 124 may calculate this loss by taking the difference of the Gram matrix values of three max pooling layers in the VGG16 model of the original and reconstructed images (i.e., the same difference used for the perceptual loss) to obtain a measure of the total variation in higher level image features. Style Loss 2: A loss similar to the effective loss, but the module 124 calculates the loss using a composite image (including the original image in the unmasked regions and the reconstructed / output image in the masked regions) instead of the reconstructed / output image used for the effective loss. - Variation loss: A measure of the transition from masked to unmasked regions in the reconstructed image.
[0096] In other embodiments, more, fewer, and / or different loss types may be used to train model 1200. At each iteration, depending on how well model 1200 reconstructs a particular input / source image (as measured based on the loss being minimized), module 124 may adjust values or parameters of model 1200 (e.g., adjust the convolution weights).
[0097] To generate synthetic “good” images from original (e.g., real) “faulty” images, model 1200 is extensively trained using good / fault-free images. In some implementations, module 124 randomly generates masks used during training (e.g., masks applied to different instances of input pair 1202). Masks may, for example, be entirely composed of lines with different widths, lengths, and positions / orientations. As a more specific example, for a 256×256 image, module 124 may randomly generate masks containing seven lines, each with a line width of 50-100 points. FIG. 13 shows two example masks 1302, 1304 of this type that may be generated by module 124. In general, masks with lines that are too narrow require very long training times, while masks with lines that are too wide result in unrealistic inpainting. In other implementations, module 124 randomly generates masks using other shapes (e.g., rectangles, circles, a mixture of shapes, etc.) and / or selects from a set of pre-designed masks.
[0098] Once model 1200 is trained in this manner, module 124 can input a defect image with a corresponding mask that obscures the defect into model 1200. FIG. 12 shows an example in which module 124 applies defect image 1210 (showing a foreign object on a syringe plunger) and defect-obscuring mask 1212 to model 1200 as input pair 1202. Trained model 1200 then reconstructs image 1210 as defect-free image 1214. Module 124 can then superimpose image 1214 onto a portion of the full container image that corresponds to the original position of input image 1210. In other implementations, module 124 can input an image of an entire container (or other object) into model 1200, and model 1200 can output a reconstructed image of the entire container (or other object).
[0099] FIG. 14 shows three exemplary sequences 1402, 1404, and 1406 in which a synthetic 256×256 image (the right side of FIG. 14 ) is generated by digitally removing defects from an actual 256×256 image (the left side of FIG. 14 ) using a partial convolution model similar to model 1200. As can be seen in the exemplary sequences 1402, 1404, and 1406, a mask is generated that can selectively obscure defects on or near the syringe plunger. Specifically, defects in the plunger itself are masked in sequence 1402, and foreign objects resting on the plunger are masked in sequences 1404 and 1406. The masks may be generated manually or by module 124 using, for example, object detection techniques. As can be seen in sequence 1406, the mask may be irregularly shaped (e.g., not symmetrical about any axis).
[0100] FIG. 15 shows another example of a 256×256 synthetic image (right side of FIG. 15 ) generated by digitally removing defects from a 256×256 real image (left side of FIG. 15 ) using a partially convolutional model similar to model 1200, along with a difference image (center of FIG. 15 ) showing how the real image was modified to arrive at the synthetic image. The difference image shows that the main change from the real image was the removal of the plunger defect, but that some noise was also added to the real image. As mentioned above, this noise can help reduce overfitting of the AVI model (neural network) during training.
[0101] FIG. 16 shows a real image 1600 of a syringe with a plunger defect and a defect-free synthetic image 1602 generated using a partial convolution model similar to model 1200. In this example, images 1600 and 1602 are both 251 × 1651 images. For this particular example, reconstruction was more efficient by first cropping a square portion of image 1600 showing the defect and generating a mask of the smaller cropped image. After reconstructing the cropped region using the partial convolution model, the reconstructed region was reinserted into original image 1600 to obtain synthetic image 1602. As can be seen in FIG. 16, synthetic image 1602 provides a realistic depiction of a defect-free syringe. Furthermore, although not easily visible to the naked eye, synthetic image 1602 contains additional noise, which can aid in the training process as described above. However, in this case, due to the cropping technique used, the added noise is not distributed throughout image 1602. In some implementations, one or more post-processing techniques may be used to ensure a more realistic transition between the reconstructed region and the surrounding region and / or to remove or minimize any artifacts. For example, after generating the composite image 1602 by reinserting the reconstructed region into the original image 1600, module 124 may add noise distributed throughout image 1602 and / or perform smoothing on image 1602.
[0102] In some implementations, module 124 also, or instead, uses deep learning-based inpainting (e.g., a partial convolution model similar to model 1200) in reverse to generate a synthetic “flawed” image from an original “good” image. In a first embodiment, this can be achieved by training a partial convolution model (e.g., model 1200) in the same manner as described above for the case of adding defects (e.g., using a good image for input pair 1202). However, to add defects, a different image is input to the trained partial convolution model. Specifically, instead of inputting a “good” image, module 124 first adds a desired defect image to the good image at the desired location. This step can use simple image processing techniques, such as simply replacing a portion of the good image with the desired defect image. Module 124 can obtain the defect image from, for example, feature image library 142. FIG. 17 shows three exemplary defect images 1700A-1700C that may be included in feature image library 142, any one of which may be used to replace a portion of the original image. Alternatively, any other suitable defect type may be used (e.g., any of the defect types described above in connection with feature image library 142 of FIG. 1, or defect types associated with other contexts such as automotive body inspection).
[0103] In some implementations, after the defect image is positioned in a desired location (e.g., with input from a user of the software tool via a graphical user interface or entirely by module 124), module 124 automatically creates a mask by setting an occlusion region to have the same size and location in the original image as the superimposed defect image. Module 124 can then input the modified original image (with the superimposed defect image) and the mask as separate inputs to a partial convolution model (e.g., model 1200).
[0104] 18 shows two exemplary sequences 1800, 1802 in which this technique is used to add defects to a 256 x 256 partial syringe image. In sequence 1800, module 124 acquires real image 1804A, superimposes desired defect image 1804B at selected (e.g., manually or randomly determined) or predetermined locations, generates mask 1804C that matches the size of real image 1804A but has occluded regions that match the size and location of superimposed defect image 1804B, and applies the modified real image and mask 1804C as separate inputs to a partial convolution model (e.g., model 1200) to generate composite image 1804D. Similarly, in sequence 1802, module 124 acquires real image 1810A, superimposes desired defect image 1810B at selected (e.g., manually or randomly determined) or predetermined locations, generates mask 1810C with occlusion regions that match the size of real image 1810A but the size and location of the superimposed defect image 1810B, and applies the modified real image and mask 1810C as separate inputs to a partial convolution model (e.g., model 1200) to generate composite image 1810D. A mask line width of 16 pt was used to generate composite images 1804D, 1810D. As can be seen in FIG. 18 , this technique blurs the masked areas with the applied defects, providing smooth transition regions with a realistic appearance. Another example is shown in FIG. 19 , where this same technique was used to augment real 251 × 1651 image 1900 to obtain composite defect image 1902.
[0105] In another implementation, module 124 uses a partial convolution model such as model 1200 to add defects to the original image, but trains the model differently to support random defect generation. In this implementation, during training, module 124 feeds each defect image (e.g., an actual defect image) into the partial convolution model to serve as a target image. The training samples are the same defect images, but with a mask that (when applied to the defect image) masks the defects. By repeating this for many defect images, module 124 trains the partial convolution model to inpaint each mask / hole region with a defect. Once the partial convolution model is trained, module 124 can apply good / defect-free images as input pairs, along with the mask at the desired defect location.
[0106] In these implementations, if multiple defect types are desired, it may be advantageous to train separate partial convolution models for different defect types. For example, module 124 may train a first partial convolution model to enhance good images by adding specks and a second partial convolution model to enhance images by adding incorrect plunger ribs, etc. This generally provides more control over defect inpainting and allows different models to be trained independently (e.g., with different hyperparameters to account for the different complexities associated with each defect type). This may also generate more “pure” defects (i.e., distinct within a single defect class), which may be useful, for example, when the composite images are used to train a computer vision system that identifies different defect classes. FIG. 20 shows three example sequences 2000, 2002, and 2004 in which this technique was used to add defects to syringe images. For each sequence, module 124 takes a real image (left side of FIG. 20), generates a mask that occludes the portion of the real image where the defect should be added (center of FIG. 20), and applies the real image and mask as separate inputs to a trained partially convolutional model (similar to model 1200) to generate a synthetic image (right side of FIG. 20). Another example is shown in FIG. 21, where this same technique was used to augment a real 251×1651 image 2100 to obtain a synthetic defect image 2102.
[0107] In some implementations, module 124 also, or instead, uses deep learning-based inpainting (e.g., a partially convolutional model similar to model 1200) to modify (e.g., move and / or change the appearance of) features shown in the original (e.g., real) image. For example, module 124 can move and / or change the appearance of a meniscus (e.g., in a syringe). In these implementations, module 124 can use either of the two techniques described above in the context of adding defects using a partially convolutional model (e.g., model 1200): (1) train a model using a “good” image as a target image, then superimpose the original image with a feature image (e.g., from feature image library 142) that exhibits the desired feature appearance / location to generate a composite image, or (2) train a model using images that exhibit the desired feature appearance / location (with a corresponding mask that obscures the feature), then mask the original image at the desired feature location to generate a composite image. An exemplary sequence 2200 for generating a composite image using the latter of these two options is shown in FIG. 22. As can be seen in FIG. 22, the mask, which may be irregularly shaped, must occlude both the portion of the original image where the relevant feature (here, the meniscus) appears and the portion of the original image where the feature will be transposed. Another example is shown in FIG. 23, where this same technique was used to expand an actual 251×1651 image 2300 (specifically, by moving the meniscus to a new position and “reshaping” the meniscus) to obtain a composite image 2302. Similar to the reconstruction shown in FIG. 16, reconstruction was more efficient by first cropping a square portion of image 2300 where the meniscus appears, and then generating a mask of the smaller cropped image. After reconstructing the cropped region using a partial convolution model, the reconstructed region was reinserted into the original image 2300 to obtain a composite image 2302.
[0108] Module 124 can also, or instead, use this technique to move / alter other features, such as the plunger (by digitally moving the plunger along the barrel) or the lyophilization vial contents (e.g., by digitally changing the fill level of the vial). In implementations where the partially convolved model is trained using target images exhibiting desired feature locations / appearances (i.e., the latter of the two techniques described above), module 124 can train and use a different model for each feature type. For a given partially convolved model, the range and variation of features (e.g., meniscus) that the model artificially generates can be adjusted by controlling variation between training samples. In general, extending features such as the meniscus to a standard state can aid in training the AVI classification model by preventing feature variation (e.g., different meniscus positions) from "confusing" the classifier while helping the classifier focus only on defects.
[0109] Inpainting using a partial convolution model can be very efficient. In the case of meniscus expansion, for example, a single base mask can be used to generate thousands of images in a few minutes, depending on the available processing power (e.g., for processing unit 110). Defect creation can be similarly efficient. In the case of defect removal, where a mask is drawn for each image to cover the defect (which can take about 1 second per image), the output can be slower (e.g., at thousands of images per hour, depending on how quickly each mask can be created). However, all of these processes are much faster and less costly than manually creating and removing defects in actual samples.
[0110] In some implementations, processing power constraints may limit the size of the image to be augmented (e.g., images of approximately 512 x 512 pixels or smaller), which may require cropping the image before augmentation and then reinserting the cropped augmented image. This may require additional time and other undesirable results (e.g., in the case of deep learning-based inpainting techniques, not achieving the benefit of adding subtle noise / variation to the entire image, not just the small / cropped portion, as described above in connection with FIG. 16). In some implementations, module 124 addresses this by using a ResNet feature extractor rather than a VGG feature extractor. Feature extractors such as these are used to calculate a loss, which is used to adjust the weights of the inpainting model during training. Module 124 may use any appropriate version of a ResNet feature extractor (e.g., ResNet50, ResNet101, ResNet152, etc.), depending on the image dimensions and the desired training speed.
[0111] Additionally, in some implementations, module 124 may apply post-processing to the synthetic images to reduce undesirable artifacts. For example, module 124 may add noise to each synthetic image, perform filtering / smoothing on each synthetic image, and / or perform fast Fourier transform (FFT) frequency spectrum analysis and manipulation on each synthetic image. Such techniques can help mitigate any artifacts and generally make the images more realistic. As another example, module 124 may pass each synthetic image through a refiner, which is trained by pairing the refiner with a discriminator. During training, both the refiner and the discriminator are fed (e.g., by module 124) with synthetic and real images. The discriminator's goal is to discriminate between real and synthetic images, and the refiner's goal is to refine the synthetic image to the point where the discriminator can no longer distinguish the synthetic image from the real image. In this way, the refiner and the discriminator are adversarial to each other, acting similar to a generative adversarial network (GAN). After multiple cycles of training, the refiner can become very adept at refining images, and module 124 can then use the trained refiner to remove artifacts from composite images that are added to training image library 140. Any of the techniques described above can also be used to process / refine composite images generated without deep learning techniques, such as composite images generated using algorithm 400 described above.
[0112] Various tests were conducted to demonstrate that generating complementary synthetic images from original images (e.g., synthetic "faulty" images versus real "good" images, or synthetic "good" images versus real "faulty" images) can significantly improve the training of AVI deep learning models (e.g., image classifiers) and guide the AVI models to accurately identify defect locations. In one such test, a ResNet50 defect classifier for syringes was trained on two training sample sets. The first training sample set consisted of 270 original images with defects and 270 original images without defects. In the second training sample set, the defect-free samples consisted of the 270 original images and 270 synthetic images (originally generated from the defective samples and with defects removed using an inpaint tool), while the defective samples consisted of the 270 original images (used to generate the synthetic defect-free images) and 270 synthetic images (generated from the 270 original defective images using an inpaint tool without a mask). In both cases, the test samples were 60 original images, a mixture of defective and non-defective images. Notably, the test samples were not independent of the training samples, as they were images from the same syringe as the training samples, differing only in rotation.
[0113] The table below summarizes the details of these training sample sets that were used to train two different AVI image classification models ("Classifier 1" and "Classifier 2").
[0114] [Table 7]
[0115] [Table 8]
[0116] Classifier 1 and Classifier 2 were each trained for eight epochs using the Adam optimizer with a learning rate of 0.0001. FIG. 24A shows Grad-CAM images 2400 and 2402 generated using Classifier 1 and Classifier 2, respectively, for a black and white stain defect. Both Classifier 1 and Classifier 2 provided 100% accuracy for the test sample used, and it can be seen from FIG. 24A that Classifier 2 provided a dramatic improvement over Classifier 1. Specifically, Classifier 2 focused on the correct region of the sample image (the plunger rib), while Classifier 1 instead focused on the meniscus region where no defect was present. Furthermore, only Classifier 1 provided the correct classification ("Defect") because, as noted above, this image was related by rotation to samples the classifier had already seen during training. Another example is shown in FIG. 24B, which shows Grad-CAM images 2410 and 2412 generated using Classifier 1 and Classifier 2, respectively, for a speck defect. Again, Classifier 2 focuses on the correct regions, while Classifier 1 focuses on the incorrect regions. This was also true for the three other defect classes tested. Thus, including 50% synthetic images in the training sample set dramatically improved the performance of the classifiers in all cases tested.
[0117] To properly train an AVI model (e.g., an image classification model), it is prudent to include quality control measures at one or more stages. This can be particularly important in the context of medicine, where patient safety must be protected by ensuring safe and reliable medicines. In some implementations, both "pre-processing" and "post-processing" quality checks are performed (e.g., by the image / library evaluation module 126). Generally, these pre-processing and post-processing quality checks can utilize various image processing techniques to analyze and / or compare information on a pixel-by-pixel basis.
[0118] Because images are typically captured under tightly controlled conditions, there are often only subtle differences between any two images from the same dataset. While measuring the variation in image parameters across an entire dataset can be laborious, a quick, visual assessment of such variation can save time (e.g., by avoiding measurement of the wrong attribute) and serve as an initial quality check on image capture conditions. Knowing this variation can be useful for two reasons. First, variation in a particular attribute (e.g., plunger position) can overwhelm the signal from the actual defect, resulting in misclassification because the algorithm may weight the varying attribute more heavily than the defect itself. Second, for image augmentation purposes, knowing the range of variation for a given attribute can be useful for constraining those attributes to that range when creating a representative composite image.
[0119] 25 shows an exemplary process 2500 for generating a visualization that can be used to quickly assess diversity in an image set. Process 2500 may be performed by image / library assessment module 126 (also referred to simply as “module 126”). In process 2500, module 126 converts image set 2502 into a set of respective numeric matrices 2504, each having exactly one matrix element for each pixel in the corresponding image from image set 2502. Module 126 then determines the maximum value across all of numeric matrices 2504 at each matrix position (i,j) and uses the maximum value to populate corresponding positions (i,j) in maximum value matrix 2506. Module 126 then converts maximum value matrix 2506 into a maximum variation composite (bitmap) image 2508. Alternatively, module 126 can avoid creating a new maximum value matrix 2506 and instead update a particular numeric matrix from set 2504 (e.g., by successively comparing each element value of that numeric matrix with the corresponding element values of all other numeric matrices 2504 and updating each time a greater value is found).
[0120] The computer system 104 can then present the resulting composite image 2508 on a display to allow for quick visualization of the variability of the dataset. FIG. 26A shows an example of such a visualization 2600. In this example, it can be seen that the plunger has moved as far to the left as point 2602. This may or may not be allowed depending on the desired constraints. The module 124 can then use point 2602, for example, as the leftmost boundary of the plunger (e.g., in creating composite images with different plunger positions). In some implementations, the module 124 determines this boundary more precisely by determining points (e.g., pixel locations) where the first derivative over successive columns exceeds a certain threshold.
[0121] Other variations of visualization 2600 are possible. For example, module 126 may determine the minimum (i.e., taking the minimum element value at each matrix position across all numeric matrices 2504) image, or the average (i.e., taking the average value at each matrix position across all numeric matrices 2504) image, etc. An exemplary average image visualization 2604 is shown in Figure 26B. In any of these implementations, this technique can be used to display variability as a quality check and / or to determine attribute / feature boundaries that the composite image must adhere to.
[0122] 27 shows an exemplary process 2700 for assessing the similarity between a composite image and a set of images. Process 2700 may be performed by image / library assessment module 126, for example, to assess composite images generated by library expansion module 124. Module 126 can use process 2700 in addition to one or more other techniques (e.g., assessing AVI model performance before and after a composite image is added to a training set). However, process 2700 is used in a more targeted manner to ensure that each composite image is not fundamentally different from the original real image.
[0123] At block 2702 of process 2700, for every image in the real image set, module 126 calculates the mean squared error (MSE) relative to every other image in the real image set. The MSE between any two images is the average of the squared differences of pixel values at all locations (e.g., at corresponding matrix element values). For example, for an i×j image, the MSE is the sum of the squared differences over all i×j pixel / element locations divided by the quantity i×j. Thus, module 126 calculates the MSE for every possible image pair in the real image set. The real image set may include all available real images or a subset of a larger set of real images.
[0124] In block 2704, module 126 determines the maximum MSE among all the MSEs calculated in block 2702 and sets an upper bound equal to the maximum MSE. This upper bound can serve, for example, as a maximum allowable amount of dissimilarity between the synthetic image and the set of real images. The lower bound is necessarily zero.
[0125] At block 2706, module 126 calculates the MSE between the synthetic image under consideration and all images in the set of real images. Thereafter, at block 2708, module 126 determines whether the largest of the MSEs calculated at block 2706 is greater than the upper boundary set at block 2704. If so, at block 2710, module 126 generates an indication of dissimilarity of the synthetic image relative to the set of real images. For example, module 126 may cause the display of an indicator that the upper boundary has been exceeded or the generation of a flag indicating that the synthetic image should not be added to the training image library 140. If the largest MSE calculated at block 2706 is not greater than the upper boundary set at block 2704, then at block 2712, module 126 does not generate an indication of dissimilarity. For example, module 126 may cause the display of an indicator that the upper boundary has not been exceeded or the generation of a flag indicating that the synthetic image should or may be added to the training image library 140.
[0126] In some implementations, process 2700 varies in one or more respects from that shown in FIG. 27 . For example, in block 2708, module 126 may instead determine whether the average of all MSEs calculated in block 2706 exceeds an upper boundary. As another example, in some implementations, module 126 generates a histogram of the MSEs calculated in block 2706 instead of (or in addition to) performing blocks 2708, 2710, or blocks 2708, 2712. An example of such a histogram 2800 is shown in FIG. 28 . The X-axis of the exemplary histogram 2800 indicates MSE, and the Y-axis indicates the number of times the MSE occurred during the comparison of the synthetic and real images. While using MSE as a quality proxy has some inherent limitations, this metric can provide a reasonable approach to complement the analysis of AVI model performance.
[0127] In some implementations, in addition to or instead of the techniques described above (e.g., process 2700), computer system 104 determines one or more other image quality metrics (e.g., to determine the similarity between a given composite image and other images, to measure the diversity of a set of images, etc.) For example, computer system 104 may use any of the techniques described in U.S. Provisional Patent Application No. 63 / 020,232 for this purpose.
[0128] 29-32 show flow diagrams of example methods corresponding to the various techniques described above. Referring first to FIG. 29, a method 2900 for generating a composite image may be performed, for example, by module 124 of FIG. 1 (e.g., when processing unit 110 executes instructions of module 124 stored in memory unit 114) by transferring features onto an original image.
[0129] At block 2902, a feature matrix is received or generated. The feature matrix is a numerical representation of a feature image that indicates a feature. A feature may be, for example, a defect associated with a container (e.g., a syringe, vial, cartridge, etc.) or the contents of the container (e.g., a fluid or lyophilized pharmaceutical), such as a crack, a chip, a stain, a foreign object, etc. Alternatively, a feature may be a defect associated with another object (e.g., a scratch or dent in the body of an automobile, a dent or crack in the siding of a house, a crack, air bubble, or impurity in a glass window, etc.). Block 2902 may include, for example, performing the defect image transformation of block 404 of FIG. 4A. In some embodiments, block 2902 includes rotating and / or resizing the feature matrix or rotating and / or resizing the image from which the feature matrix is derived (e.g., as described above in connection with FIG. 4A for the more specific case in which the “feature” is a defect). If the feature image is rotated and / or resized, this step is performed before generating the feature matrix to ensure that the feature matrix reflects the rotation. If block 2902 includes rotating the feature matrix or feature image, method 2900 may include rotating the feature matrix or feature image by an amount based on both (1) the rotation of the features depicted in the feature image and (2) the desired rotation of the features depicted in the feature image. Method 2900 may include determining this "desired" rotation based, for example, on the location of the region to which the features are to be transferred. Block 2902 may also, or instead, include resizing the feature matrix or feature image.
[0130] In block 2904, a surrogate region matrix is received or generated. The surrogate region matrix is a numerical representation of the region in the original image to which features will be transferred / transposed. Block 2904 may be similar to, for example, block 410 of FIG. 4A.
[0131] In block 2906, the feature matrix is normalized for portions of the feature matrix that do not represent the indicated feature. Block 2906 may include, for example, block 412 of FIG. 4A.
[0132] A composite image is generated based on the surrogate region matrix and the normalized feature matrix at block 2908. Block 2908 may include, for example, blocks 414, 416, 418, 420, 422, and 424 of FIG.
[0133] It is understood that the blocks of method 2900 do not have to occur in the exact order shown. For example, blocks 2906 and 2908 may occur in parallel, block 2904 may occur before block 2902, etc.
[0134] Referring now to FIG. 30, a method 3000 for generating a composite image may be performed, for example, by module 124 of FIG. 1 (e.g., when processing unit 110 executes instructions of module 124 stored in memory unit 114) by removing defects shown in the original image.
[0135] At block 3002, a portion of the original image that shows a defect is masked. The mask may be applied, for example, automatically (e.g., by first detecting the defect using object detection) or in response to user input identifying an appropriate masked region.
[0136] At block 3004, a correspondence index is calculated, which reflects pixel statistics that indicate the correspondence between portions of the original image adjacent to the masked portion and other portions of the original image.
[0137] In block 3006, the correspondence indices calculated in block 3004 are used to fill the masked portions of the original image with defect-free image portions, for example, the masked portions may be filled / painted in a manner that attempts to mimic other patterns in the original image.
[0138] In block 3008, a neural network is trained for automated visual inspection using the synthetic image (e.g., multiple other real images and the synthetic image). The AVI neural network may be, for example, an image classification neural network or an object detection (e.g., convolutional) neural network.
[0139] It is understood that the blocks of the method 3000 do not have to occur in the exact order shown.
[0140] Referring now to FIG. 31 , a method 3100 for generating a composite image by removing or modifying features shown in the original image or by adding features shown in the original image may be performed, for example, by module 124 of FIG. 1 (e.g., when processing unit 110 executes instructions of module 124 stored in memory unit 114).
[0141] In block 3102, a partially convolutional model (e.g., similar to model 1200) is trained. The partially convolutional model includes an encoder with a series of convolutional layers and a decoder with a series of transposed convolutional layers. Block 3102 includes, for each image in the training image set, applying the training image and corresponding mask as separate inputs to the partially convolutional model.
[0142] At block 3104, a composite image is generated. Block 3104 includes, for each of the original images, applying the original image (or a modified version of the original image) and the corresponding mask as separate inputs to a trained partially convolutional model. The original image may be first modified, for example, by superimposing a cropped image of the feature to be added (e.g., a defect) before applying the modified original image and the corresponding mask as inputs to the trained partially convolutional model.
[0143] In block 3106, a neural network for automated visual inspection is trained using the synthetic image (and possibly also using the original image). The AVI neural network may be, for example, an image classification neural network or an object detection (e.g., convolutional) neural network.
[0144] It is understood that the blocks of the method 3100 do not have to occur in the exact order shown.
[0145] Referring now to FIG. 32, for example, a method 3200 for evaluating synthetic images that may be used in a training image library may be performed by module 124 of FIG. 1 (e.g., when processing unit 110 executes instructions of module 124 stored in memory unit 114).
[0146] In block 3202, a measure of the difference between (1) each image in an image set (e.g., real images) and (2) each other image in the image set is calculated based on the pixel values of the images. Block 3202 may be similar to, for example, block 2702 of FIG. 27.
[0147] In block 3204, a threshold difference value (e.g., the "upper boundary" in FIG. 27) is generated based on the metric calculated in block 3202. Block 3204 may be similar to, for example, block 2704 in FIG.
[0148] Various operations are repeated for each of the composite images in block 3206. In particular, in block 3208, a composite image index is calculated based on pixel values of the composite image, and in block 3210, acceptability of the composite image is determined based on the composite image index and the threshold difference value. For example, block 3208 may be similar to block 2706 of FIG. 27, and block 3210 may include block 2708 of FIG. 27 and either block 2710 or block 2712. In some implementations, block 3206 includes one or more manual steps (e.g., manually determining acceptability based on a displayed histogram similar to histogram 2800 shown in FIG. 28).
[0149] It is understood that the blocks of the method 3200 do not have to occur in the exact order shown.
[0150] Although the systems, methods, devices, and components thereof have been described in terms of exemplary embodiments, they are not limited to these exemplary embodiments. The detailed description is to be construed as an example only and does not describe every possible embodiment of the invention, as describing every possible embodiment would be impractical, if not impossible. Many alternative embodiments can be implemented using either current technology or technology developed after the filing date of this patent, and still fall within the scope of the claims that define the invention.
[0151] Those skilled in the art will understand that various modifications, variations and combinations can be made to the above-described embodiments without departing from the scope of the present invention, and that such modifications, variations and combinations should be construed as being within the scope of the present invention.
Claims
1. 1. A method for generating a synthetic image by transferring features onto an original image, comprising: receiving or generating a feature matrix that is a numerical representation of the feature image indicative of the feature, each element of the feature matrix corresponding to a different pixel of the feature image; receiving or generating a surrogate region matrix that is a numerical representation of a region in the original image to which the feature is to be transferred, where each element of the surrogate region matrix corresponds to a different pixel in the original image; normalizing the feature matrix for portions of the feature matrix that do not represent the feature; generating the composite image based on (i) the surrogate region matrix and (ii) the normalized feature matrix; A method comprising:
2. the original image is an image of a container, the feature is a defect related to the container or the contents of the container; The method of claim 1.
3. the container is a syringe, the feature is a defect associated with the barrel of the syringe, the plunger of the syringe, the needle shield of the syringe, or the fluid within the syringe; The method of claim 2.
4. the container is a vial, the feature is a defect associated with the wall of the vial, the cap of the vial, the crimp of the vial, or the fluid or lyophilized cake within the vial; The method of claim 2.
5. The method of any one of claims 1 to 4, wherein normalizing the feature matrix comprises normalizing the feature matrix row-wise or column-wise.
6. Normalizing the feature matrix row by row or column by column includes, for each row or column of the feature matrix, generating a feature row histogram of element values of the row or column of the feature matrix; The method of claim 5 , comprising:
7. Normalizing the feature matrix row by row or column by column includes, for each row or column of the feature matrix, identifying peaks in the feature row histogram that correspond to portions of the rows or columns of the feature matrix that do not represent the features; For each element of the row or column of the feature matrix, subtracting the center value of the peak portion from the value of the element; The method of claim 6 further comprising:
8. 8. The method of claim 7, wherein subtracting the central value of the peak portion from the value of the element comprises: (i) subtracting an average value of all values in the row or column corresponding to the peak portion from (ii) the value of the element.
9. For each row or column of the proxy region matrix, generating a surrogate region row histogram; Identifying a peak portion of the surrogate region row histogram; determining a range of values representing the width of the peak portion of the surrogate region row histogram; Further comprising: generating the composite image includes generating the composite image based on (i) the numerical ranges for each row or column of the feature matrix and (ii) the normalized feature matrix. The method according to any one of claims 1 to 8.
10. generating the composite image, for each row or column of the feature matrix, For each element of the row or column of the feature matrix, determining whether the element of the feature matrix has a value within the numerical range; modifying an original image matrix, which is a numerical representation of the original image, by either (i) preserving the original value of a corresponding element of the original image matrix if the element of the feature matrix has a value within the numerical range, or (ii) setting the corresponding element in the original image matrix equal to the sum of the original value and the value of the element in the feature matrix if the value of the element of the feature matrix is not within the numerical range; 10. The method of claim 9, comprising:
11. The method of claim 10 , wherein generating the composite image comprises converting the modified original image matrix into a bitmap image.
12. The method of any one of claims 1 to 11, wherein receiving or generating the feature matrix comprises rotating the feature matrix or the feature image.
13. 13. The method of claim 12, wherein rotating the feature matrix or the feature image comprises rotating the feature matrix or the feature image by an amount based on (i) a rotation of the features depicted in the feature image and (ii) a desired rotation of the features depicted in the feature image.
14. determining the desired rotation based on the location of the region to which the features are to be transferred; The method of claim 13 further comprising:
15. The method of any one of claims 1 to 14, wherein receiving or generating the feature matrix comprises resizing the feature matrix or the feature image.
16. repeating the method for each of a plurality of features corresponding to different features in a feature library; The method of any one of claims 1 to 15, further comprising:
17. generating a plurality of composite images by repeating the method for each of a plurality of original images; The method of any one of claims 1 to 16, further comprising:
18. training a neural network for automated visual inspection using the plurality of composite images and the plurality of original images; 20. The method of claim 17, further comprising:
19. inspecting a plurality of images for the indicated defect using the trained neural network; 20. The method of claim 18, further comprising:
20. 1. A system comprising: one or more processors; One or more non-transitory computer-readable media storing instructions that, when executed by the one or more processors, cause the system to: receiving or generating a feature matrix that is a numerical representation of a feature image, where each element of the feature matrix corresponds to a different pixel of the feature image; receiving or generating a surrogate region matrix that is a numerical representation of a region in an original image to which the feature is to be transferred, each element of the surrogate region matrix corresponding to a different pixel of the original image; normalizing the feature matrix for portions of the feature matrix that do not represent the feature; generating the synthetic image based on (i) the surrogate region matrix and (ii) the normalized feature matrix; one or more non-transitory computer-readable media; A system comprising:
21. 21. The system of claim 20, wherein normalizing the feature matrix comprises normalizing the feature matrix row-wise or column-wise.
22. Normalizing the feature matrix row by row or column by column includes, for each row or column of the feature matrix, generating a feature row histogram of element values of the row or column of the feature matrix; 22. The system of claim 21, comprising:
23. Normalizing the feature matrix row by row or column by column includes, for each row or column of the feature matrix, identifying peaks in the feature row histogram that correspond to portions of the rows or columns of the feature matrix that do not represent the features; For each element of the row or column of the feature matrix, subtracting the center value of the peak portion from the value of the element; 23. The system of claim 22, further comprising:
24. 24. The system of claim 23, wherein subtracting the central value of the peak portion from the value of the element comprises: (i) subtracting an average value of all values in the row or column corresponding to the peak portion from (ii) the value of the element.
25. The instructions further include: For each row or column of the proxy region matrix, Generate a surrogate region row histogram; Identifying a peak portion of the surrogate region row histogram; determining a range of values representing a width of the peak portion of the surrogate region row histogram; generating the composite image includes generating the composite image based on (i) the numerical ranges for each row or column of the feature matrix and (ii) the normalized feature matrix. A system according to any one of claims 20 to 24.
26. generating the composite image, For each row or column of the feature matrix, For each element of the row or column of the feature matrix, determining whether the element of the feature matrix has a value within the numerical range; modifying an original image matrix, which is a numerical representation of the original image, by either (i) preserving the original value of a corresponding element of the original image matrix if the element of the feature matrix has a value within the numerical range, or (ii) setting the corresponding element of the original image matrix equal to the sum of the original value and the value of the element in the feature matrix if the value of the element of the feature matrix is not within the numerical range; converting the modified original image matrix into a bitmap image; 26. The system of claim 25, comprising:
27. The system of any one of claims 20 to 26, wherein receiving or generating the feature matrix comprises rotating the feature matrix or the feature image.
28. 28. The system of claim 27, wherein rotating the feature matrix or the feature image comprises rotating the feature matrix or the feature image by an amount based on (i) a rotation of the features depicted in the feature image and (ii) a desired rotation of the features depicted in the feature image.
29. The instructions further include: determining the desired rotation based on the location of the region to which the features are to be transferred; 29. The system of claim 28.
30. The system of any one of claims 20 to 29, wherein receiving or generating the feature matrix comprises resizing the feature matrix or the feature image.
31. 1. A method for generating a composite image by removing defects shown in an original image, comprising: masking the portion of the original image that shows the defect; generating a composite image, the composite image comprising, at least in part, calculating a correspondence index indicating a correspondence between (i) portions of the original image adjacent to the masked portion and (ii) other portions of the original image; and filling the masked portion of the original image with a defect-free image portion using the calculated correspondence indices. generating a composite image by using the synthetic image to train a neural network for automated visual inspection; A method comprising:
32. Calculating the correspondence index calculating a first statistic for each of a plurality of adjacent portions of the original image adjacent the masked portion; calculating second statistics for each of a plurality of portions of the original image that are outside both (i) the masked portion and (ii) the plurality of adjacent portions; calculating the correspondence indicator based on the first statistic and the second statistic; 32. The method of claim 31 , comprising:
33. the original image is an image of a container, the defect is associated with the container or the contents of the container; 33. The method of claim 31 or 32.
34. the container is a syringe, the defect is associated with the barrel of the syringe, the plunger of the syringe, the needle shield of the syringe, or the fluid within the syringe; 34. The method of claim 33.
35. the container is a vial, the defect is associated with the wall of the vial, the cap of the vial, the crimp of the vial, or the fluid or lyophilized cake within the vial; 34. The method of claim 33.
36. generating a plurality of composite images by repeating the method for each of a plurality of original images; Further comprising: training the neural network for automated visual inspection includes training the neural network using the plurality of composite images.
36. The method according to any one of claims 31 to 35.
37. 37. The method of claim 36, wherein training the neural network for the automated visual inspection comprises training the neural network using the plurality of composite images and the plurality of original images.
38. inspecting a plurality of images using the trained neural network; 38. The method of any one of claims 31 to 37, further comprising:
39. 1. A system comprising: one or more processors; One or more non-transitory computer-readable media storing instructions that, when executed by the one or more processors, cause the system to: masking portions of the original image that show defects; generating a composite image, the composite image comprising, at least in part, calculating a correspondence index indicating a correspondence between (i) portions of the original image adjacent to the masked portion and (ii) other portions of the original image; and filling the masked portion of the original image with a defect-free image portion using the calculated correspondence indices. generating a composite image by using the synthetic image to train a neural network for automated visual inspection; one or more non-transitory computer-readable media for causing the A system comprising:
40. Calculating the correspondence index calculating a first statistic for each of a plurality of adjacent portions of the original image adjacent the masked portion; calculating second statistics for each of a plurality of portions of the original image that are outside both (i) the masked portion and (ii) the plurality of adjacent portions; calculating the correspondence indicator based on the first statistic and the second statistic; 40. The system of claim 39, comprising:
41. the original image is an image of a container, the defect is associated with the container or the contents of the container; 41. A system according to claim 39 or 40.
42. the container is a syringe, the defect is associated with the barrel of the syringe, the plunger of the syringe, the needle shield of the syringe, or the fluid within the syringe; 42. The system of claim 41.
43. the container is a vial, the defect is associated with the wall of the vial, the cap of the vial, the crimp of the vial, or the fluid or lyophilized cake within the vial; 42. The system of claim 41.
44. 1. A method of generating a composite image by removing or modifying features shown in an original image or by adding features shown in said original image, comprising: For each of a plurality of training images, training a partially convolutional model including (i) an encoder having a series of convolutional layers and (ii) a decoder having a series of transposed convolutional layers by applying at least the training image and a corresponding training mask as a separate input to the partially convolutional model; generating the composite image by applying, for each of the original images, at least (i) either the original image or a modified version of the original image, and (ii) a corresponding mask as separate inputs to the trained partially convolutional model; training a neural network for automated visual inspection using the synthetic image; A method comprising:
45. 45. The method of claim 44, wherein the method generates the composite image by removing defects shown in the original image.
46. generating the composite image includes, for each of the original images, applying the original image and the corresponding mask as separate inputs to the trained partially convolutional model; the original image is an image showing a corresponding defect, the corresponding mask, when applied to the original image, obscures the corresponding defect; 46. The method of claim 45.
47. 47. The method of claim 46, wherein the training images, the original image, and the composite image depict a container.
48. 48. The method of claim 47, wherein the container is a syringe.
49. 49. The method of claim 48, wherein the corresponding defect is a syringe barrel defect, a syringe plunger defect, or a defect related to syringe contents.
50. 48. The method of claim 47, wherein the container is a vial.
51. 51. The method of claim 50, wherein the corresponding defect is a vial wall defect, a vial cap defect, a vial crimp defect, or a defect related to the vial contents.
52. 45. The method of claim 44, wherein the method generates the composite image by adding indicated defects to the original image or modifying indicated features in the original image.
53. generating the composite image comprises, for each of the original images: modifying the original image by replacing certain portions of the original image with feature images; applying the modified original image and the corresponding mask as separate inputs to the trained partially convolutional model; Including, the corresponding mask, when applied to the modified original image, obscures the particular portion of the original image; 53. The method of claim 52.
54. 54. The method of claim 53, wherein generating the composite image comprises automatically generating, for each of the original images, the corresponding mask to correspond to the particular portion of the original image.
55. 55. The method of claim 53 or 54, wherein the training images, the original image, and the composite image show a container.
56. 56. The method of claim 55, wherein the container is a syringe.
57. the method generating the composite image by adding the indicated defects to the original image; The characteristic image is an image of a defect related to a syringe barrel defect, a syringe plunger defect, or a syringe content.
57. The method of claim 56.
58. 56. The method of claim 55, wherein the container is a vial.
59. the method generating the composite image by adding the indicated defects to the original image; the characteristic image is a defect image related to a vial wall defect, a vial cap defect, a vial crimp defect, or a vial content; 59. The method of claim 58.
60. the method generating the composite image by modifying features shown in the original image; the characteristic image is an image of a meniscus; 56. The method of any one of claims 53 to 55.
61. the corresponding training mask, when applied to the training image, obscures the corresponding features shown in the training image; generating the composite image includes, for each of the original images, applying the original image and the corresponding mask as separate inputs to the trained partially convolutional model; 53. The method of claim 52.
62. 62. The method of claim 61 , wherein the training images, the original image, and the composite image depict a container.
63. 63. The method of claim 62, wherein the container is a syringe.
64. the method generating the composite image by adding the indicated defects to the original image; the corresponding feature is a syringe barrel defect, a syringe plunger defect, or a defect related to the syringe contents; 64. The method of claim 63.
65. 63. The method of claim 62, wherein the container is a vial.
66. the method generating the composite image by adding the indicated defects to the original image; the corresponding feature is a vial wall defect, a vial cap defect, a vial crimp defect, or a defect related to the vial contents; 64. The method of claim 63.
67. training the partially convolutional model includes training the partially convolutional model to add a particular type of defect; the corresponding feature corresponds to the particular type of defect; 67. The method of any one of claims 61 to 66.
68. the method generating the composite image by modifying features shown in the original image; the corresponding feature is a meniscus; 62. The method of claim 61.
69. 69. The method of any one of claims 44 to 68, wherein training the partially convolutional model comprises minimizing a plurality of losses, the plurality of losses comprising at least an effectiveness loss, a hole loss, a perceptual loss, a style loss, and a variation loss.
70. 1. A system comprising: one or more processors; One or more non-transitory computer-readable media storing instructions that, when executed by the one or more processors, cause the system to: For each of a plurality of training images, training a partially convolutional model including (i) an encoder having a series of convolutional layers and (ii) a decoder having a series of transposed convolutional layers by applying at least the training image and a corresponding training mask as a separate input to the partially convolutional model; generating a plurality of composite images by applying, for each of a plurality of original images, at least (i) either the original image or a modified version of the original image, and (ii) a corresponding mask as separate inputs to the trained partially convolutional model; training a neural network for automated visual inspection using the synthetic image; one or more non-transitory computer-readable media; A system comprising:
71. 71. The system of claim 70, wherein the system generates the composite image by removing defects shown in the original image.
72. generating the composite image includes, for each of the original images, applying the original image and the corresponding mask as separate inputs to the trained partially convolutional model; the original image is an image showing a corresponding defect, the corresponding mask, when applied to the original image, obscures the corresponding defect; 72. The system of claim 71.
73. 73. The system of claim 72, wherein the training images, the original image, and the composite image depict a container.
74. 71. The system of claim 70, wherein the system generates the composite image by adding indicated defects to the original image or modifying features indicated in the original image.
75. generating the composite image comprises, for each of the original images: modifying the original image by replacing certain portions of the original image with feature images; applying the modified original image and the corresponding mask as separate inputs to the trained partially convolutional model; Including, the corresponding mask, when applied to the modified original image, obscures the particular portion of the original image; 75. The system of claim 74.
76. 76. The system of claim 75, wherein the training images, the original image, and the composite image show a container.
77. the system generates the composite image by adding the indicated defects to the original image; the feature image is an image of a defect; 77. A system according to claim 75 or 76.
78. the system generates the composite image by modifying features shown in the original image; the characteristic image is an image of a meniscus; 77. A system according to claim 75 or 76.
79. the corresponding training mask, when applied to the training image, obscures the corresponding features shown in the training image; generating the composite image includes, for each of the original images, applying the original image and the corresponding mask as separate inputs to the trained partially convolutional model; 75. The system of claim 74.
80. 80. The system of claim 79, wherein the training images, the original image, and the composite image show a container.
81. the system generates the composite image by adding the indicated defects to the original image; the corresponding features are defects shown in the training images; 81. A system according to claim 79 or 80.
82. training the partially convolutional model includes training the partially convolutional model to add a particular type of defect; the corresponding feature corresponds to the particular type of defect; 82. The system of claim 81.
83. the system generates the composite image by modifying features shown in the original image; the corresponding feature is a meniscus; 81. A system according to claim 79 or 80.
84. 84. The system of any one of claims 70 to 83, wherein training the partially convolutional model comprises minimizing a plurality of losses, the plurality of losses comprising at least an effectiveness loss, a hole loss, a perceptual loss, a style loss, and a variation loss.
85. 1. A method for evaluating synthetic images used in a training image library, comprising: calculating, based on pixel values of the set of images, a metric indicative of the difference between (i) each image in the set of images and (ii) each other image in the set of images; generating a threshold difference value based on the calculated index; For each of the plurality of composite images, calculating a composite image measure based on pixel values of the composite image; determining acceptability of the composite image based on the composite image indicator and the threshold difference value; A method comprising:
86. for each of the plurality of composite images, either adding the composite image to the training image library or omitting the composite image from the training image library based on the acceptability of the composite image; 86. The method of claim 85, further comprising:
87. training a neural network for automated visual inspection using the training image library after adding or omitting each of the composite images; 87. The method of claim 86, further comprising:
88. inspecting a plurality of container images for indicated defects using the trained neural network; 88. The method of claim 87, further comprising:
89. calculating the measure indicative of the difference between each image and each other image; calculating a mean squared error between (i) the pixel values of each image in the set of images and (ii) the pixel values of each other image in the set of images; 89. The method of any one of claims 85 to 88, comprising:
90. generating the threshold difference value, determining a maximum mean squared error from among the mean squared errors calculated for image pairs formed from the set of images; 90. The method of claim 89, comprising:
91. Calculating the composite image indices comprises: calculating a mean squared error for a plurality of composite images, wherein the mean squared error for each composite image is the mean squared error between (i) the pixel values of the composite image and (ii) the pixel values of each image in the set of images; 91. The method of claim 89 or 90, comprising:
92. 92. The method of claim 91, wherein determining the acceptability of the composite image is based on (i) a maximum value of mean squared errors of the composite image and (ii) the maximum mean squared error.
93. 92. The method of claim 91, wherein determining the acceptability of the composite image is based on (i) an average of the mean squared errors of the composite image and (ii) the maximum mean squared error.
94. determining the acceptability of the composite image; presenting on a display a histogram representing the composite image mean square error of said composite image; 94. The method of any one of claims 91 to 93, comprising:
95. 95. The method of any one of claims 85 to 94, wherein the set of images is a set of digitally unaltered container images and the composite image is a digitally altered container image.
96. 1. A system comprising: one or more processors; One or more non-transitory computer-readable media storing instructions that, when executed by the one or more processors, cause the system to: calculating, based on pixel values of the set of images, a metric indicative of the difference between (i) each image in the set of images and (ii) each other image in the set of images; generating a threshold difference value based on the calculated index; For each of the plurality of composite images, calculating a composite image index based on pixel values of the composite image; determining acceptability of the composite image based on the composite image indicator and the threshold difference value; one or more non-transitory computer-readable media; A system comprising:
97. calculating the measure indicative of the difference between each image and each other image; calculating a mean squared error between (i) the pixel values of each image in the set of images and (ii) the pixel values of each other image in the set of images; 97. The system of claim 96, comprising:
98. generating the threshold difference value, determining a maximum mean squared error from among the mean squared errors calculated for image pairs formed from said image set; 98. The system of claim 97, comprising:
99. Calculating the composite image indices comprises: calculating a mean squared error for a plurality of composite images, wherein the mean squared error for each composite image is the mean squared error between (i) the pixel values of the composite image and (ii) the pixel values of each image in the set of images; 99. The system of claim 97 or 98, comprising:
100. determining the acceptability of the composite image; the maximum of the mean squared error and the maximum mean squared error of the composite image; or the mean squared error of the composite image and the average of the maximum mean squared error; 100. The system of claim 99, based on
101. The instructions further include: presenting on a display a histogram representing the mean square error of the composite image of the composite image; 101. A system according to claim 99 or 100.
102. 102. The system of any one of claims 96 to 101, wherein the set of images is a set of digitally unaltered container images and the composite image is a digitally altered container image.