Visual inspection model and its learning apparatus, method and program, and visual inspection system
Patent Information
- Application Number
- JP2025026041
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2026-09-01
AI Technical Summary
【0011】 本発明によると、符号化照明下の対象物のワンショットカラー画像に基づくイメージング外観検査のための外観検査モデルおよびその学習において、光源数をあらかじめ指定した数に削減して最適化を行うことができる。さらに、外観検査システムにおいて光源数を削減することができ、コストを低減することができる。
Smart Images

Figure 2026139396000001_ABST
Abstract
Description
[Technical Field]
[0001] This invention relates to imaging-based visual inspection, a visual inspection technique based on images, and more particularly to a machine learning model suitable for imaging-based visual inspection, as well as its learning and application. [Background technology]
[0002] Identifying the type of material, such as metal or plastic, or the material itself, such as iron or aluminum, and its condition, including cracks, scratches, corrosion, rust, and contamination, is crucial for recycling, visual inspection, and understanding texture. While visual inspection methods include eddy current, ultrasound, and lasers, image-based visual inspection (imaging visual inspection) is superior to other methods because it is non-destructive and non-contact, has high identification accuracy, and offers a fast response time.
[0003] Conventional imaging-based visual inspection methods acquire images by sequentially illuminating multiple light sources, resulting in a large number of images. Even when using encoded illumination that simultaneously illuminates multiple light sources, at least two images are required, making real-time inspection based on single-shot images difficult. Furthermore, the need for multiple light sources presents cost challenges.
[0004] To address these challenges, the inventors have developed a method that simultaneously optimizes the illumination environment, represented by a 1x1 convolutional kernel, and feature extraction, represented by a nonlinear function, using a Convolutional Neural Network (CNN) for one-shot color images of objects captured under multi-wavelength, multi-directional coded illumination (see, for example, Non-Patent Literature 1). Furthermore, by imposing non-negative and sparse constraints on the intensity of the light sources represented by the 1x1 convolutional kernel, high-precision identification is possible even with one-shot color images under a small number of light sources. [Prior art documents] [Non-patent literature]
[0005] [Non-Patent Document 1] Wang Chao, Kawahara Ryo, Okabe Takahiro, "Simultaneous Optimization Network for Illumination Environment and Feature Extraction for Surface Material Identification," Information Processing Society of Japan Research Report, Vol.2022-CVIM-228, No.29, January 2022. [Overview of the Initiative] [Problems that the invention aims to solve]
[0006] In the technique described above, which performs convolution operations using a 1x1 convolution kernel on color images under coded illumination, the sparse constraint allows for a reduction in the number of required light sources while maintaining high-precision discrimination capabilities. However, this number cannot be specified in advance; rather, the number of light sources is determined during the optimization process. Furthermore, the simultaneous optimization of the illumination environment and feature extraction tends to require enormous computational resources and time.
[0007] Therefore, the present invention aims to provide a machine learning model and its training that can optimize imaging visual inspection based on a one-shot color image of an object under encoded illumination so that the number of light sources can be reduced to a predetermined number, and further to provide a visual inspection system with a reduced number of light sources. [Means for solving the problem]
[0008] According to one aspect of the present invention, an appearance inspection model is provided for classifying pixel by pixel a color image of an object illuminated simultaneously with multi-wavelength, multi-directional light from multiple light sources, wherein the model includes weights corresponding to the illumination intensity of each of the multiple light sources, the weights are subject to non-negative constraints and sparse constraints, and the computer functions as a channel mixing MLP that mixes the feature quantities of the pixels separately for RGB using the weights, a token mixing MLP that mixes the feature quantities mixed by the channel mixing MLP among the RGB components, and a class classifier that classifies the pixels based on the feature quantities mixed by the token mixing MLP.
[0009] According to another aspect of the present invention, there is provided an apparatus for training the visual inspection model described above, comprising: a training data input unit that acquires RGB values of pixels located at the same position in a plurality of color images of an object when the plurality of light sources illuminate one by one at unit intensity, and inputs a pixel value vector composed of these RGB values as a feature quantity of the pixel to the channel mixing MLP; a parameter optimization unit that optimizes parameters of the visual inspection model based on an error between an output result of the visual inspection model under training and a correct label such that weights of the channel mixing MLP do not become negative; and a sparsity adjustment unit that sets a specified number, which is fewer than the number of the plurality of light sources, among the weights of the channel mixing MLP to positive values during training of the visual inspection model. There are also provided a visual inspection model training apparatus, and a corresponding visual inspection model training method and program.
[0010] According to still another aspect of the present invention, there is provided a system for inspecting an appearance of an object, comprising: an encoded illumination device having a plurality of light sources, which simultaneously illuminates the object with multi-wavelength and multi-directional light from the plurality of light sources; a camera that captures a color image of the object simultaneously illuminated with the multi-wavelength and multi-directional light; and a visual inspection device that classifies the pixels by using the trained visual inspection model according to claim 1, wherein RGB values of each pixel of the color image are input to the channel mixing MLP as a linear combination value of a feature quantity of the pixel and a weight of the channel mixing MLP in the visual inspection model to classify the pixel, wherein the plurality of light sources in the encoded illumination device are limited to the number of weights made positive by the sparsity constraint of the channel mixing MLP in the trained visual inspection model, and illuminate at an intensity corresponding to the weights. A visual inspection system characterized by the above is provided. Effects of the Invention
[0011] According to the present invention, in a visual inspection model for imaging visual inspection based on a one-shot color image of an object under encoded illumination and training thereof, optimization can be performed by reducing the number of light sources to a predetermined number. Furthermore, the number of light sources can be reduced in the visual inspection system, thereby reducing costs. [Brief Description of Drawings]
[0012] [Figure 1] FIG. 1 is a schematic diagram of an appearance inspection model according to an embodiment of the present invention. [Figure 2] FIG. 2 is a diagram explaining the relationship between an image under coded illumination and an image under a single light source. [Figure 3] FIG. 3 is a schematic diagram representing dynamic sparse learning for coded illumination. [Figure 4] FIG. 4 is a schematic configuration diagram of a computer device used for learning and inference of an appearance inspection model. [Figure 5] FIG. 5 is a flowchart of an appearance inspection model learning method according to an embodiment of the present invention. [Figure 6] FIG. 6 is a functional block diagram of an appearance inspection model learning device according to an embodiment of the present invention. [Figure 7] FIG. 7 is a schematic configuration diagram of an appearance inspection system according to an embodiment of the present invention. [Mode for Carrying Out the Invention]
[0013] Hereinafter, embodiments of the present invention will be described in detail with appropriate reference to the drawings. However, excessive detailed description may be omitted. For example, detailed description of already well-known matters and redundant description of substantially the same configuration may be omitted. This is to avoid unnecessary redundancy in the following description and to facilitate understanding for those skilled in the art. The inventor provides the accompanying drawings and the following description for those skilled in the art to fully understand the present invention, and does not intend to limit the subject matter recited in the claims by these. In addition, dimensions, thicknesses, detailed shapes of each member and the like drawn in the drawings may differ from actual ones.
[0014] [Appearance Inspection Model] Figure 1 is a schematic diagram of a visual inspection model according to one embodiment of the present invention. The visual inspection model 10 according to this embodiment classifies the material and condition such as scratches and rust for each pixel of a color image of an object captured under encoded illumination. Encoded illumination refers to multi-wavelength, multi-directional light produced by simultaneously illuminating with multiple light sources at various intensities. The object to be classified is assumed to be an opaque flat surface with no coating on its surface.
[0015] The visual inspection model 10 incorporates the techniques of an MLP (Multilayer Perceptron) mixer, a type of neural network used for image processing, and includes a channel mixing MLP 11, a token mixing MLP 12, and a classifier 13. Note that diagrams illustrating processes such as transposition (swapping the dimensions of features before and after the token mixing MLP 12) and layer normalization (normalizing the inputs to the channel mixing MLP 11 and token mixing MLP 12) are omitted.
[0016] The channel mixing MLP11 and token mixing MLP12 themselves are constructed as fully connected layers + activation function + fully connected layers, similar to typical MLP mixers. The class classifier 13 includes global average pooling, fully connected layers, and Softmax, similar to class classifiers used in CNNs. The probability of each class is output from the final Softmax, and the class with the highest probability becomes the predicted class.
[0017] In a typical MLP mixer, the spatial relationships of image data are first learned using a token mixing MLP, and then the relationships between features for each channel are learned using a channel mixing MLP. In contrast, the appearance inspection model 10 first performs feature mixing between channels using a channel mixing MLP 11, and then performs feature mixing between tokens using a token mixing MLP 12.
[0018] FIG. 2 is a diagram explaining the relationship between an image under coded illumination and an image under a single light source. Assume that there are three light sources A, B, and C having different wavelengths and directions, the unit intensity (maximum intensity) of each light source is "1", and the coded illumination performs simultaneous illumination such that light source A has an illumination intensity of "0.1", light source B has an illumination intensity of "0.5", and light source C has an illumination intensity of "0.2". In this case, from the principle of light superposition, an image captured under coded illumination is obtained by weighting the respective illumination intensities and adding together the images obtained when each light source independently illuminates at unit intensity. That is, the image under coded illumination is obtained by linearly combining the images under single light source illumination, which serve as inputs of a perceptron, with the illumination intensities of the respective light sources, which serve as the connection strengths, i.e., weights, of the perceptron, and this corresponds to mixing feature values between channels. For this reason, in the appearance inspection model 10, the channel mixing MLP 11 is arranged before the token mixing MLP 12.
[0019] The reflection characteristic at an arbitrary point on the surface of an object illuminated by a given light source depends on the direction of incident light (θ i , φ i ), the direction of reflected light (θ o , φ o ) and the wavelength λ, and is described by a five-dimensional spectral BRDF (Bidirectional Reflectance Distribution Function) f(θ i , φ i , θ o , φ o , λ). When a plurality of light sources arranged in D different directions with C different colors (spectral intensities) are used, spectral BRDF feature values of the object surface can be sampled in both the incident angle direction (θ i , φ i ) and the wavelength λ.
[0020] Let a vector obtained by arranging pixel values of rgb bands observed at an arbitrary point of a color image of an object captured by illuminating L (L=C×D) light sources one by one at unit intensity be a spectral BRDF feature quantity {x r , x g , x b}, then the pixel value of the rgb band observed at each point on the surface of the object under coded illumination by L light sources (I r,I g ,I b ) ┬ It is given by equation (1) below. TIFF2026139396000002.tif22158 Here, w is an encoded illumination vector representing the illumination intensity of each of the L light sources.
[0021] Equation (1) can be viewed as a weighted summation of features between different channels, which is similar to the mixing of features between channels in a channel mixing MLP in an MLP mixer. Therefore, in the visual inspection model 10, the parameter weights in the channel mixing MLP 11 are associated with the illumination intensity of each light source in the encoded illumination, thereby linking the spectral BRDF features with the encoded illumination.
[0022] Let N=3, then X∈R N×L Let be the input matrix for spectral BRDF features. L-dimensional vector x r , x g , x b This consists of the RGB values observed at any point on the surface of the object when each of the L light sources illuminates it at a unit intensity. The appearance inspection model 10 uses a 3L-dimensional feature space (x r ,x g ,x b ) ┬ The spectral BRDF features, represented as follows, are classified. The computation process in channel mixing MLP11 is expressed as follows (2). TIFF2026139396000003.tif11156 Here, W c1 ∈R L×H and W c2 ∈R H×L is the weight matrix, b c1 and b c2 σ is the bias vector, H is the hidden dimension, and σ is the activation function such as GELU. H=1 and W c1 By setting it to ≥0, the weight matrix W c1 This becomes the encoded illumination vector w.
[0023] The computational process in the token mixing MLP12 for feature extraction is expressed as follows (3). TIFF2026139396000004.tif11155 Here, W t1 ∈R L×H ', W t2 ∈R H ′ ×L , b t1 and b t2 H' is a learnable parameter, and H' is the hidden dimension.
[0024] The final classification calculation process in the classifier 30 is expressed as shown in equation (4). TIFF2026139396000005.tif10152
[0025] [Dynamic sparse learning of coded illumination] Figure 3 is a schematic diagram representing the dynamic sparse learning of encoded illumination. In the visual inspection model 10, in order to reduce the number of light sources required for encoded illumination, a non-negative constraint and a sparse constraint are imposed on the encoded illumination vector w as shown in equations (5) and (6). TIFF2026139396000006.tif14156 Here, k is the number of active light sources. Active light sources are the light sources that are illuminating.
[0026] Initialize k values of w with random positive values, and set the rest to zero. Then, perform dynamic sparse learning, repeatedly updating the sparse pattern of w during the learning process.
[0027] In the dynamic sparse learning of encoded illumination, the sparse pattern is effectively adjusted by combining weight growth based on the gradient of the weights of the channel mixing MLP11 with weight reduction based on the magnitude of the weights of the channel mixing MLP11. If the weights are reduced too much and fall below k, the weights of the channel mixing MLP11 are grown anew based on the magnitude of the gradient. This allows the visual inspection model 10 to explore diverse combinations of light source direction and wavelength by activating weights that can significantly reduce losses. Weight reduction is performed by removing relatively small weights, thereby eliminating less influential light sources and maintaining the target sparseness level.
[0028] Dynamic sparse learning is performed by repeatedly executing the following steps. 1. Forward processing: Calculate the model output using the current parameters, including w. The input data uses the RGB values of pixels at the same position in multiple color images of an object when illuminated by multiple light sources, one by one, at unit intensity.
[0029] 2. Loss calculation: Cross-entropy loss L based on model output results and ground truth labels. CE Calculate.
[0030] 3. Reverse processing: Calculate the gradient ∇L for all parameters.
[0031] 4. Gradient Mask: Apply a sparse constraint mask m to the gradient of w as shown in equation (7). In the formula TIFF2026139396000007.tif9157, the circled dot operator represents element-wise multiplication. This ensures that in channel mixing MLP11, only gradients corresponding to active weights are used for updates. Active weights are positive weights.
[0032] 5. Parameter update: Update w with the learning rate η using an optimizer such as Adam, as shown in equation (8) below. TIFF2026139396000008.tif9154
[0033] 6. Enforcing Non-Negativity: We enforce the non-negativity constraint by applying equation (9). This reflects the physical reality that the intensity of a light source cannot be negative. TIFF2026139396000009.tif9157
[0034] 7. Dynamic Spatial Adjustment: Perform the following dynamic spatial adjustment after repeating steps 1 through 7 above T times.
[0035] (a) Weight pruning: Identify weights with relatively small values and set them to zero so that there are exactly k non-zero elements in w.
[0036] (b) Weight Growth: Weights with a value of zero that have a relatively large gradient are activated, and weights that have been removed and are no longer numbered below k are replenished. This allows the visual inspection model 10 to adaptively select promising light sources based on the current loss situation. If the absolute value of the gradient does not reach the required number, weights are filled in using random positions so that the number of newly grown weights equals k.
[0037] (c) Mask update: After weight reduction and weight growth, the sparse constraint mask m is updated to reflect the new sparse pattern, so that it has exactly k active elements. This mask indicates whether each light source in the coded illumination is active (1) or inactive (0), and is reapplied to the gradient in subsequent updates to maintain the sparse constraint.
[0038] This dynamic sparse learning enables efficient exploration of the encoded illumination space, promotes diversity in selected light sources, and improves classification performance while adhering to hardware constraints. By adjusting the sparse pattern based on both the magnitude of the weights and the gradient information, the visual inspection model 10 can satisfy the sparse constraints on encoded illumination while maintaining sufficient representational power.
[0039] [Simultaneous optimization with the classification network] The overall optimization problem can be expressed as follows (10). TIFF2026139396000010.tif11155 Here, θ includes all learnable parameters in the visual inspection model 10.
[0040] By simultaneously optimizing the parameters of the visual inspection model 10 and the encoded illumination vector in this way, it is ensured that the learned encoded illumination is specially tailored for the classification task while satisfying physical and hardware constraints.
[0041] [Model execution environment]
[0042] The visual inspection model 10 is trained using multiple color images of an object, each illuminated at a unit intensity by multiple light sources constituting the coded illumination. At this time, the weights of the channel mixing MLP 11 are associated with the illumination intensity of each of the multiple light sources constituting the coded illumination. A non-negative constraint is imposed on the weights of the channel mixing MLP 11 to prevent the illumination intensity from becoming negative, and a sparse constraint is imposed to reduce the number of light sources. In inference using the trained visual inspection model 10, the color images of the object under coded illumination with the reduced number of light sources due to the sparse constraint are input as the values of the hidden nodes of the channel mixing MLP 11, and the object is classified pixel by pixel.
[0043] Figure 4 is a schematic diagram of the computer system used for training and inference of the visual inspection model. The computer system 100 includes a processor 101, memory 102, storage 103, and interface 104. These elements are interconnected via a bus or the like. The computer system 100 may consist of a single computer system, or it may be configured as a distributed system with multiple servers connected via a network.
[0044] The processor 101 is the element that reads the program and parameters of the visual inspection model 10 stored in the storage 103, expands them into memory 102, and performs execution and learning processing. The processor 101 may be a general-purpose CPU, GPU, or dedicated hardware (ASIC, FPGA, etc.).
[0045] Memory 102 is an element for storing temporary data such as intermediate calculation results, model parameter expansion data, input data, and output data that are necessary when the processor 101 executes a program or performs learning and inference processing of the visual inspection model 10.
[0046] Storage 103 is a storage device for long-term storage of various data, including the program for training and inference processing the visual inspection model 10 and the visual inspection model 10 itself. The visual inspection model 10 has its model parameters and model structure stored in storage 103, and the processor 101 reads these as needed, expands them into memory 102, and performs training and inference processing. As storage 103, an HDD, SSD, or other non-volatile storage device can be used.
[0047] Interface 104 is an element for sending and receiving data, issuing operation instructions, and displaying results with users and external systems. Interface 104 includes a GUI, a communication interface, and interfaces for input / output devices such as keyboards and mice.
[0048] [Model Learning] Figure 5 is a flowchart of a visual inspection model learning method according to one embodiment of the present invention. Figure 6 is a functional block diagram of a visual inspection model learning device according to one embodiment of the present invention. The visual inspection model learning device 20 according to this embodiment is a device that performs the above-mentioned model learning in the flow shown in Figure 5, and can be realized with a computer device as shown in Figure 4.
[0049] As an example, the visual inspection model learning device 20 includes a learning data input unit 21, a parameter optimization unit 22, and a sparse adjustment unit 23. The learning data input unit 21 performs the "1. Forward processing" described above, reading image data of the object from the storage 103 when multiple light sources constituting encoded illumination illuminate it one by one at a unit intensity, obtaining the RGB values of pixels at the same position in each image, and inputting a pixel value vector consisting of these RGB values as a pixel feature quantity to the channel mixing MLP 11 (S1). This pixel value vector corresponds to the spectral BRDF feature quantity X in equation (2). The weight initialization unit 231 in the sparse adjustment unit 23 initializes the weight W of the channel mixing MLP 11. c1 A specified number of k elements are randomly selected and kept, while the rest are set to zero (S1).
[0050] The parameter optimization unit 22 optimizes the parameters of the visual inspection model 10 so that the weights of the channel mixing MLP 11 do not become negative, based on the error between the output results of the visual inspection model 10 during training and the correct labels (S2). More specifically, the parameter optimization unit 22 includes a loss calculation unit 221, a gradient calculation unit 222, a gradient mask unit 223, a parameter update unit 224, and a non-negative force unit 225.
[0051] The loss calculation unit 221 performs the "2. Loss Calculation" described above, and calculates the cross-entropy loss based on the error between the output result of the visual inspection model 10 and the correct label. The gradient calculation unit 222 performs the "3. Gradient Calculation" described above, and calculates the gradient of the parameters of the visual inspection model 10 based on the cross-entropy loss. The gradient mask unit 223 performs the "4. Gradient Mask" described above, and applies the sparse constraint mask held in the storage 103 to the gradient of the weights of the channel mixing MLP 11 among the gradients. The parameter update unit 224 performs the "5. Parameter Update" described above, and updates the parameters of the visual inspection model 10 based on the gradient after applying the sparse constraint mask. The non-negativity enforcement unit 225 performs the "6. Non-negativity Enforcement" described above, and sets any negative values in the weights of the channel mixing MLP 11 after the parameter update to zero.
[0052] The sparse adjustment unit 23 performs the "7. Dynamic sparse adjustment" described above, and sets a specified number of weights of the channel mixing MLP 11 that are less than the number of light sources to positive values during the training of the visual inspection model 10. More specifically, the sparse adjustment unit 23 comprises a weight initialization unit 231, a weight reduction unit 232, a weight growth unit 233, and a mask update unit 234.
[0053] The weight initialization unit 231 randomly selects a specified number of weights from the channel mixing MLP 11 and keeps them while setting the others to zero at the start of training the visual inspection model 10. The weight reduction unit 232 performs the above-mentioned "(a) weight reduction," and if the number of positive weights in the channel mixing MLP 11 is greater than the specified number (> specified number in S3), it sets the relatively small weights to zero to maintain the number of positive weights at the specified number (S4). The weight growth unit 233 performs the above-mentioned "(b) weight growth," and if the number of positive weights in the channel mixing MLP 11 is less than the specified number (< specified number in S3), it activates the weights with relatively large gradients to maintain the number of positive weights at the specified number (S5). The mask update unit 234 performs the "(c) mask update" described above, and updates the sparse constraint mask held in the storage 103 to reflect the weight pattern of the channel mixing MLP 11 adjusted by the weight reduction unit 232 and the weight growth unit 233 (S6).
[0054] After sparse adjustment, if the loss of the visual inspection model 10 converges or the number of training iterations reaches the specified number (YES in S7), training of the visual inspection model 10 is terminated. Otherwise, the process returns to step S2, and the parameter optimization and dynamic sparse training of the coded illumination are repeated.
[0055] [Examples of applications of pre-trained models] Figure 7 is a schematic diagram of a visual inspection system according to one embodiment of the present invention. The visual inspection system 30 according to this embodiment uses a trained visual inspection model 10 to inspect the appearance of an object 200 from a one-shot color image of the object 200. As an example, the visual inspection system 30 includes an encoded illumination device 31, a camera 32, and a visual inspection device 33.
[0056] The coding illumination device 31 has multiple light sources 311 and is a device that simultaneously illuminates an object 200 with multi-wavelength, multi-directional light from multiple light sources 311. For example, the coding illumination device 31 can be realized with an LED-based multispectrum dome. In a multispectrum dome, several light sources 311 made of LEDs that emit light of various wavelengths come together to form a light source cluster 310, and multiple light source clusters 310 are fixed at multiple locations on the frame in various directions.
[0057] The full set of light sources is shown within the frame of the figure. The full set of light sources is an illumination device used when capturing color images to be used as training data for the appearance inspection model 10, and consists of a total of L (L = C × D) light sources arranged in D different directions with C different colors (spectral intensities). In contrast, the coded illumination device 31 is limited to a smaller number of light sources 311, that is, the number of weights that have been made positive by the sparse constraint of the channel mixing MLP 11 in the trained appearance inspection model 10, and illuminates with an intensity corresponding to those weights.
[0058] Camera 32 is a device that captures a color image of an object 200 that is simultaneously illuminated with multi-wavelength, multi-directional light from an encoded illumination device 31. The appearance inspection device 33 is a device that inspects the appearance of an object 200 by classifying the one-shot color image acquired from camera 32 pixel by pixel using a trained appearance inspection model 10, and can be realized with a computer device as shown in Figure 4. Specifically, the appearance inspection device 33 calculates the RGB value of each pixel of the color image acquired from camera 32 as a linear combination value of the pixel feature quantity and the weight of the channel mixing MLP 11 in the appearance inspection model 10 (XW of equation (2)). c1 The data is input to the channel mixing MLP 11, processed through the token mixing MLP 12 and the class classifier 13, and the class classification of each pixel is performed.
[0059] <Effects> According to this embodiment, in the visual inspection model 10 for imaging visual inspection based on a one-shot color image of an object under encoded illumination, and in its training, the number of light sources can be reduced to a predetermined number for optimization. Furthermore, the number of light sources in the visual inspection system 30 can be reduced, thereby lowering costs.
[0060] ≪Variations≫ The sparse constraints imposed on the channel mixing MLP 11 of the visual inspection model 10 are not limited to specifying the number of light sources. For example, when training the visual inspection model 10 based on training data acquired using multiple cameras, multiple shooting conditions, multiple sensors, etc., sparse constraints may be imposed to limit the number of cameras, shooting conditions, the number of sensors, etc.
[0061] As described above, embodiments have been explained as examples of the technology in the present invention. For this purpose, accompanying drawings and a detailed description have been provided. Therefore, among the components described in the accompanying drawings and detailed description, there may be not only components that are essential for solving the problem, but also components that are not essential for solving the problem, in order to illustrate the above technology. For this reason, the mere fact that these non-essential components are described in the accompanying drawings and detailed description should not be immediately assumed to be essential. Furthermore, since the above embodiments are for the purpose of illustrating the technology in the present invention, various changes, substitutions, additions, omissions, etc., can be made within the scope of the claims or equivalents. [Explanation of Symbols]
[0062] 10 Visual Inspection Models 11-Channel Mixing MLP 12 Token Mixing MLP 13 Classifiers 20. Visual Inspection Model Learning Device 21. Learning Data Input Section 22 Parameter Optimization Unit 221 Loss calculation section 222 Gradient Calculation Unit 223 Gradient Mask Section 224 Parameter Optimization Unit 225 Non-negative compulsion section 23 Sparse adjustment section 231 Weight Initialization Section 232 Weight reduction unit 233 Weight growth section 234 Mask Update Section 30 Visual Inspection System 31 Coded lighting device 311 Light source 32 cameras 33 Visual Inspection Device 200 Objects
Claims
1. A visual inspection model for classifying pixel by pixel a color image of an object that has been simultaneously illuminated by multi-wavelength, multi-directional light from multiple light sources, A channel mixing MLP that includes weights corresponding to the illumination intensity of each of the multiple light sources, with non-negative and sparse constraints imposed on the weights, and mixes the feature quantities of the pixels separately for RGB using the weights. A token mixing MLP that mixes the feature quantities mixed by the channel mixing MLP between RGB, and An appearance inspection model characterized by using a computer as a classifier that classifies pixels based on the feature quantities mixed by the aforementioned token mixing MLP.
2. The appearance inspection model according to claim 1, wherein the sparse constraint specifies the number of the plurality of light sources to be illuminated.
3. A device for training the visual inspection model described in claim 1, A learning data input unit that acquires the RGB values of pixels at the same position in multiple color images of an object when each of the multiple light sources illuminates it at a unit intensity, and inputs a pixel value vector consisting of these RGB values as the feature quantity of the pixel to the channel mixing MLP. A parameter optimization unit optimizes the parameters of the visual inspection model so that the weights of the channel mixing MLP do not become negative, based on the error between the output result of the visual inspection model during training and the correct label. An appearance inspection model learning device characterized by comprising: a sparse adjustment unit that sets a specified number of the weights of the channel mixing MLP, which are less than the number of light sources, to a positive value during the learning of the appearance inspection model.
4. The parameter optimization unit, A loss calculation unit that calculates the cross-entropy loss based on the error between the output result of the visual inspection model and the correct label, A gradient calculation unit that calculates the gradient of the parameters of the visual inspection model based on the cross-entropy loss, A gradient mask unit that applies a sparse constraint mask to the gradient of the weight of the channel mixing MLP among the gradients, in accordance with the sparse constraints imposed on the channel mixing MLP, A parameter update unit updates the parameters of the visual inspection model based on the gradient after applying the sparse constraint mask, The appearance inspection model learning device according to claim 3, further comprising a non-negative force unit that sets negative weights in the channel mixing MLP after the parameter update to zero.
5. The sparse adjustment unit, A weight initialization unit at the start of training the aforementioned visual inspection model randomly selects a specified number of weights from the channel mixing MLP and keeps them, while setting the others to zero. If the number of positive weights in the channel mixing MLP exceeds the specified number, a weight reduction unit reduces the relatively small weights to zero to maintain the number of positive weights at the specified number. If the number of positive weights in the channel mixing MLP is less than the specified number, a weight growth unit activates weights with relatively large gradients to maintain the number of positive weights at the specified number. The appearance inspection model learning device according to claim 4, further comprising: a mask update unit that updates the sparse constraint mask to reflect the weight pattern of the channel mixing MLP adjusted by the weight reduction unit and the weight growth unit.
6. The sparse adjustment unit, A weight initialization unit at the start of training the aforementioned visual inspection model randomly selects a specified number of weights from the channel mixing MLP and keeps them, while setting the others to zero. If the number of positive weights in the channel mixing MLP exceeds the specified number, a weight reduction unit reduces the relatively small weights to zero to maintain the number of positive weights at the specified number. If the number of positive weights in the channel mixing MLP is less than the specified number, a weight growth unit activates weights with relatively large gradients to maintain the number of positive weights at the specified number. The appearance inspection model learning device according to claim 3, further comprising: a mask update unit that updates a sparse constraint mask in accordance with the sparse constraints imposed on the channel mixing MLP, which is referenced when the parameter update unit updates the parameters of the appearance inspection model, so as to reflect the weight pattern of the channel mixing MLP adjusted by the weight reduction unit and the weight growth unit.
7. A method for training the visual inspection model described in claim 1, Computers A learning data input step involves obtaining the RGB values of pixels at the same position in multiple color images of an object when each of the multiple light sources illuminates it at a unit intensity, and inputting a pixel value vector consisting of these RGB values as the feature quantity of the pixel into the channel mixing MLP. A parameter optimization step that optimizes the parameters of the visual inspection model so that the weights of the channel mixing MLP do not become negative, based on the error between the output result of the visual inspection model during training and the correct label, A method for learning an appearance inspection model, characterized by performing a sparse adjustment step during the learning of the appearance inspection model, in which a specified number of the weights of the channel mixing MLP that are less than the number of light sources are set to positive values.
8. The parameter optimization step described above is A loss calculation step that calculates the cross-entropy loss based on the error between the output result of the visual inspection model and the correct label, A gradient calculation step of calculating the gradient of the parameters of the visual inspection model based on the cross-entropy loss, A gradient mask step in which a sparse constraint mask is applied to the gradient of the weight of the channel mixing MLP among the gradients, in accordance with the sparse constraints imposed on the channel mixing MLP, A parameter update step in which the parameters of the visual inspection model are updated based on the gradient after applying the sparse constraint mask, The method for learning an appearance inspection model according to claim 7, comprising a non-negative forcing step of setting negative weights in the channel mixing MLP after the parameter update to zero.
9. The aforementioned sparse adjustment step, A weight initialization step at the start of training the aforementioned visual inspection model includes randomly selecting and retaining the specified number of weights from the channel mixing MLP and setting the others to zero, If the number of positive weights in the channel mixing MLP exceeds the specified number, a weight reduction step is performed to set the relatively small weights to zero to maintain the number of positive weights at the specified number. If the number of positive weights in the channel mixing MLP is less than the specified number, a weight growth step is performed to activate weights with relatively large gradients to maintain the number of positive weights at the specified number, A method for learning an appearance inspection model according to claim 8, comprising: a mask update step of updating the sparse constraint mask to reflect the weight pattern of the channel mixing MLP adjusted by the weight reduction step and the weight growth step.
10. The aforementioned sparse adjustment step, A weight initialization step at the start of training the aforementioned visual inspection model includes randomly selecting and retaining the specified number of weights from the channel mixing MLP and setting the others to zero, If the number of positive weights in the channel mixing MLP exceeds the specified number, a weight reduction step is performed to set the relatively small weights to zero to maintain the number of positive weights at the specified number. If the number of positive weights in the channel mixing MLP is less than the specified number, a weight growth step is performed to activate weights with relatively large gradients to maintain the number of positive weights at the specified number, A method for learning an appearance inspection model according to claim 7, comprising: a mask update step, which updates a sparse constraint mask that conforms to the sparse constraints imposed on the channel mixing MLP, and which is referenced when updating the parameters of the appearance inspection model in the parameter update step, so as to reflect the pattern of weights of the channel mixing MLP adjusted by the weight reduction step and the weight growth step.
11. A program for training the visual inspection model described in claim 1, A learning data input unit acquires the RGB values of pixels at the same position in multiple color images of an object when each of the multiple light sources illuminates it at a unit intensity, and inputs a pixel value vector consisting of these RGB values as the feature quantity of the pixel to the channel mixing MLP. A parameter optimization unit optimizes the parameters of the visual inspection model so that the weights of the channel mixing MLP do not become negative, based on the error between the output result of the visual inspection model during training and the correct label, and An appearance inspection model learning program characterized by causing the computer to function as a sparse adjustment unit that sets a specified number of weights in the channel mixing MLP that are less than the number of light sources to a positive value during the learning of the appearance inspection model.
12. The parameter optimization unit, A loss calculation unit that calculates the cross-entropy loss based on the error between the output result of the visual inspection model and the correct label, A gradient calculation unit that calculates the gradient of the parameters of the visual inspection model based on the cross-entropy loss, A gradient mask unit that applies a sparse constraint mask to the gradient of the weight of the channel mixing MLP among the gradients, in accordance with the sparse constraints imposed on the channel mixing MLP, A parameter update unit updates the parameters of the visual inspection model based on the gradient after applying the sparse constraint mask, The visual inspection model learning program according to claim 11, further comprising a non-negative forcing unit that sets negative weights in the channel mixing MLP after the parameter update to zero.
13. The sparse adjustment unit, A weight initialization unit at the start of training the aforementioned visual inspection model randomly selects a specified number of weights from the channel mixing MLP and keeps them, while setting the others to zero. If the number of positive weights in the channel mixing MLP exceeds the specified number, a weight reduction unit reduces the relatively small weights to zero to maintain the number of positive weights at the specified number. If the number of positive weights in the channel mixing MLP is less than the specified number, a weight growth unit activates weights with relatively large gradients to maintain the number of positive weights at the specified number. The appearance inspection model learning program according to claim 12, further comprising: a mask update unit that updates the sparse constraint mask to reflect the weight pattern of the channel mixing MLP adjusted by the weight reduction unit and the weight growth unit.
14. The sparse adjustment unit, A weight initialization unit at the start of training the aforementioned visual inspection model randomly selects a specified number of weights from the channel mixing MLP and keeps them, while setting the others to zero. If the number of positive weights in the channel mixing MLP exceeds the specified number, a weight reduction unit reduces the relatively small weights to zero to maintain the number of positive weights at the specified number. If the number of positive weights in the channel mixing MLP is less than the specified number, a weight growth unit activates weights with relatively large gradients to maintain the number of positive weights at the specified number. The appearance inspection model learning device according to claim 11, comprising: a mask update unit that updates a sparse constraint mask in accordance with the sparse constraints imposed on the channel mixing MLP, which is referenced when the parameter update unit updates the parameters of the appearance inspection model, so as to reflect the weight pattern of the channel mixing MLP adjusted by the weight reduction unit and the weight growth unit.
15. A system for inspecting the appearance of an object, An encoding illumination device having multiple light sources, which simultaneously illuminates the object with multi-wavelength, multi-directional light from the multiple light sources, A camera that captures a color image of the object illuminated simultaneously with multi-wavelength, multi-directional light, The appearance inspection device includes a visual inspection device that uses a pre-trained visual inspection model according to claim 1 to input the RGB values of each pixel in the color image into the channel mixing MLP as a linear combination of the feature quantity of the pixel and the weight of the channel mixing MLP in the visual inspection model, thereby classifying the pixels. An appearance inspection system characterized in that the plurality of light sources in the encoded illumination device are limited to the number of weights that have been made positive by the sparse constraint of the channel mixing MLP in the learned appearance inspection model, and illuminate with an intensity corresponding to the weights.