Surface inspection device and surface inspection method

The surface inspection device employs a two-stage deep learning model with unsupervised and supervised learning to minimize manual data preparation, enhancing efficiency and accuracy in surface inspection across various manufacturing sites.

JP7807718B2Active Publication Date: 2026-01-28NIPPON STEEL CORPORATION
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2025184860
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-10-31
Publication Date
2026-01-28
Estimated Expiration
2045-10-31

AI Technical Summary

Technical Problem

Building a conventional deep learning model for surface inspection requires a large amount of manual training data, which is labor-intensive.

Method used

A surface inspection device and method that utilizes an upstream task model with a base model for unsupervised learning and a downstream task model for supervised learning, reducing the need for manual data preparation by using a combination of unsupervised and supervised data in a two-stage approach.

Benefits of technology

Reduces the effort and cost of building a deep learning model while maintaining high accuracy in surface inspection, allowing for versatile deployment across multiple manufacturing sites with reduced supervised data requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007807718000001
    Figure 0007807718000001
  • Figure 0007807718000002
    Figure 0007807718000002
  • Figure 0007807718000003
    Figure 0007807718000003
Patent Text Reader

Abstract

To provide a technique for reducing cost required for constructing a deep learning model and performing highly accurate inspection by the deep learning model.SOLUTION: A surface inspection device that inspects a surface of an object, the surface inspection device including an upstream task model including a base model and a downstream task model including a deep learning model, wherein in a learning phase, the upstream task model is learned by inputting a first learning image, and the downstream task model is learned by supervised learning by inputting a feature amount obtained by inputting a second learning image to the upstream task model and a label indicating a state of a surface of the object corresponding to the feature amount. In the inference phase, the upstream task model outputs the feature amount corresponding to the inspection image by inputting the inspection image, and the downstream task model outputs the state of the surface of the target object by inputting the feature amount corresponding to the inspection image.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a surface inspection device and a surface inspection method. [Background technology]

[0002] In recent years, many industrial fields have been actively using classifiers trained by machine learning methods to extract experience and knowledge from huge amounts of data and lead to automation. In particular, in the manufacturing industry, many machine learning models, such as deep learning models, have been introduced, and efforts are underway to improve the efficiency of inspections that previously relied on human labor. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent No. 6624963 [Patent Document 2] Japanese Patent Publication No. 2022-015575 [Patent Document 3] Special Publication No. 2024-541040 Summary of the Invention [Problem to be solved by the invention]

[0004] However, building a conventional deep learning model requires a huge amount of manual training data, which is a set of images and human-annotated correct answers, which is labor-intensive.

[0005] The present disclosure has been made in consideration of the above-mentioned circumstances, and aims to provide a technology that reduces the cost required to build a deep learning model and performs highly accurate inspections using a deep learning model. [Means for solving the problem]

[0006] In order to solve the above problem, according to one aspect of the present disclosure, there is provided a surface inspection device for inspecting the surface of an object, the surface inspection device including: an image acquisition unit that acquires captured images obtained by capturing images of the surface of the object; an upstream task model consisting of a base model that is a deep learning neural network intended to extract features of the input captured images; and a downstream task model consisting of a deep learning model, wherein in a learning phase in which the upstream task model and the downstream task model are optimized, the image acquisition unit acquires first learning captured images and second learning images that are captured images obtained by capturing images of the object, and the upstream task model learns by inputting the first learning images; The downstream task model is trained by inputting feature amounts obtained by inputting the second training image into the upstream task model and labels indicating the surface condition of the object corresponding to the feature amounts, and in an inference phase in which inference is made using the upstream task model and the downstream task model, the image acquisition unit acquires an inspection image which is an image obtained by imaging the object to be inspected, the upstream task model receives the inspection image as an input and outputs feature amounts corresponding to the inspection image, and the downstream task model receives the feature amounts corresponding to the inspection image as an input and outputs the surface condition of the object to be inspected, thereby providing a surface inspection device. In order to solve the above-described problems, according to another aspect of the present disclosure, there is provided a surface inspection device for inspecting the surface of an object, the surface inspection device including: an image acquisition unit that acquires captured images obtained by capturing images of the surface of the object; an upstream task model consisting of a base model that is a deep learning neural network intended to extract features of the input captured images; and a downstream task model that is a deep learning model, wherein in a learning phase in which the upstream task model and the downstream task model are optimized, the image acquisition unit acquires first learning captured images and second learning images that are captured images obtained by capturing images of the object, the upstream task model learns by inputting the first learning images, and the downstream task model learns by inputting feature amounts obtained by inputting the second learning images into the upstream task model and labels indicating the surface condition of the object that correspond to the feature amounts. In order to solve the above problem, according to yet another aspect of the present disclosure, there is provided a surface inspection device for inspecting a surface of an object, the surface inspection device including: an image acquisition unit that acquires captured images obtained by capturing an image of the surface of the object; and a model acquisition unit that acquires an upstream task model consisting of a base model that is a deep learning neural network intended to extract features of the input captured images; and a downstream task model consisting of a deep learning model, wherein the model acquisition unit acquires first learning captured images and second learning images that are captured images obtained by capturing an image of the object, and inputs the first learning images to acquire the upstream task model that has been trained and the second learning images. In an inference phase in which inference is performed using the upstream task model and the downstream task model, the image acquisition unit acquires an inspection image, which is an image obtained by capturing an image of the object to be inspected, and the upstream task model receives the inspection image and outputs the feature values ​​corresponding to the inspection image, and the downstream task model receives the feature values ​​corresponding to the inspection image and outputs the surface condition of the object to be inspected. The model for the upstream task and the model for the downstream task may be replaced by a model obtained by compacting the model for the upstream task and the model for the downstream task using a distillation or quantization technique in a deep learning model. In order to solve the above-described problems, according to one aspect of the present disclosure, there is provided a surface inspection method for inspecting the surface of an object, using a surface inspection device having an image acquisition unit that acquires captured images obtained by capturing an image of the surface of the object, an upstream task model consisting of a base model that is a deep learning neural network intended to extract features of the input captured images, and a downstream task model consisting of a deep learning model, and in a learning phase in which the upstream task model and the downstream task model are optimized, the method includes: an image acquisition step for acquiring first and second captured images for learning, which are captured images obtained by capturing images of the object, using the image acquisition unit; an upstream task model learning step for inputting the first learning image into the upstream task model to perform learning; The present invention provides a surface inspection method, comprising: a downstream task model learning step in which learning is performed by inputting feature amounts obtained by inputting a learning image into the upstream task model and labels indicating the surface state of the object corresponding to the feature amounts; an inspection image acquisition step in which, in an inference phase in which inference is performed using the upstream task model and the downstream task model, an inspection image acquisition step in which an inspection image is acquired by imaging the object to be inspected using the image acquisition unit; an upstream task model inference step in which the inspection image is input using the upstream task model and the feature amounts corresponding to the inspection image are output; and a downstream task model inference step in which the downstream task model is input using the downstream task model and the feature amounts corresponding to the inspection image are output. [Effects of the Invention]

[0007] According to the present disclosure, it is possible to reduce the load required to build a deep learning model and perform highly accurate inspection using the deep learning model. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is a conceptual diagram illustrating a conventional method for building a deep learning model for surface inspection. [Figure 2] FIG. 1 is a conceptual diagram illustrating a method for building a deep learning model for surface inspection in the present disclosure. [Figure 3] 1 is a schematic block diagram showing a configuration of a surface inspection device according to an embodiment of the present disclosure. [Figure 4A] 10 is a flowchart illustrating an operation in a learning phase of the surface inspection apparatus according to the embodiment of the present disclosure. [Figure 4B] 10 is a flowchart illustrating an operation of an inference phase of a surface inspection apparatus according to an embodiment of the present disclosure. [Figure 5] 1 is a graph showing the results of an example. DETAILED DESCRIPTION OF THE INVENTION

[0009] Figure 1 is a conceptual diagram showing a conventional method for constructing a deep learning model for surface inspection. Conventionally, when inspecting the surface condition (surface texture) of an object, including defects, using images, a deep learning model 400 capable of estimating the surface texture shown in the image is constructed by inputting a large amount of supervised data 220 into the deep learning model 400. More specifically, the supervised data 220 is composed of an image of the surface of an object and information (labels) indicating the correct type of surface texture obtained by an operator or inspector confirming the surface texture of the object in the image. When an image of the surface of an object included in the supervised data 220 is input, the deep learning model 400 performs a weighting calculation on the image based on weighting coefficients assigned to multiple layers and outputs an inference result. The deep learning model 400 then compares the output inference result with the label included in the supervised data 220 and performs an optimization calculation to optimize the weighting coefficients so that the output inference result matches the label. By repeating this optimization calculation, the deep learning model 400 learns and is constructed so that it can infer the type of surface texture shown in the input image. In such a method, in order to build a learning model 400 with enough accuracy to be used in the field of surface inspection, a large amount of supervised data 220 (pairs of images and labels) of several thousand to several tens of thousands may be required. Generally, supervised data (supervised data 220) is created manually, and collecting a large amount of supervised data requires a great deal of effort.

[0010] Therefore, the present inventors focused on a "base model" with the aim of reducing the load required for preparing supervised data 220. FIG. 2 is a conceptual diagram illustrating a method for constructing a deep learning model for surface inspection according to the present disclosure. The present inventors came up with the idea of ​​combining the trained base model 12 in the upstream stage with the inspection model 13 in the downstream stage, which is a deep learning model trained using conventionally known supervised data 220. This allows features of the inspection image to be extracted using the base model 12, which does not require supervised data 220 during training (i.e., less effort is required to construct the supervised data 220). This allows features of the inspection image to be extracted using the base model 12, which does not require supervised data 220 during training (i.e., less effort is required to construct the supervised data 220). Therefore, it was thought that the model that performs training using the supervised data 220 (supervised learning) could be limited to the inspection model 13. In other words, instead of training the entire process related to surface inspection using only one deep learning model as in the past, we thought that by dividing the task into two, an upstream task 110 that trains the base model 12 and a downstream task 120 that trains the inspection model 13, we could reduce the effort required for training (the effort required to obtain supervised data 220) without compromising inspection accuracy and create a defect discrimination model for surface inspection.

[0011] The surface inspection device 1 according to the present disclosure will be described below with reference to the drawings. In the following description, components having substantially the same functional configurations will be assigned the same reference numerals, and redundant description may be omitted.

[0012] (Configuration of surface inspection device 1) The surface inspection device 1 according to the present disclosure is a device used to inspect the surface of an object, that is, the surface inspection device 1 is a device that inspects the surface texture of the surface of an object. 3 is a schematic block diagram showing the configuration of a surface inspection device 1 according to an embodiment. As shown in FIG. 3, the surface inspection device 1 has an image acquisition unit 11, an upstream task model 12, and a downstream task model 13. The objects handled by the surface inspection device 1 are not particularly limited as long as they are objects that can be inspected when inspecting surface properties, but examples include steel products such as steel plates, slabs, billets, structural steel, and steel pipes in the steelmaking industry.

[0013] The image acquisition unit 11 is a functional unit that acquires a captured image obtained by capturing an image of the surface of an object. That is, the image acquisition unit 11 is configured to include a camera, and is installed at a position opposite to the inspection position (a position overlooking the conveyance device) that is set in advance on the surface of the object so that the camera can capture an image of the inspection position on the surface of the object to inspect the surface texture of the object. The image acquisition unit 11 generates and acquires a captured image, for example, of an object being conveyed by a conveyance device (not shown), by capturing an image of the inspection position on the surface of the object. The captured image acquired by the image acquisition unit 11 is used for surface inspection of the object, as described below. As described above, the image acquiring unit 11 may be provided with its own camera and may acquire the captured image by generating the captured image by itself, but it may also acquire (obtain) a captured image similar to the one captured by the image acquiring unit 11 itself via a network from outside the surface inspection device 1. For example, the image acquiring unit 11 may acquire, as the captured image, an image captured at a manufacturing base (manufacturing line) different from the manufacturing base (manufacturing line) that manufactures the object to be inspected.

[0014] The image acquisition unit 11 also functions as a functional unit that acquires first learning images and second learning images, which are captured images of an object, in a learning phase (a learning phase in which the upstream task model and the downstream task model are optimized), which will be described later. That is, the image acquisition unit 11 acquires captured images of an object (first learning images), which can be used for learning, as will be described later. Similarly to the first learning images, the image acquisition unit 11 can also acquire second learning images, which can be used for learning, as will be described later. The first training images are images used to train the upstream task model 12 in the training phase, which will be described later. Images used as the first training images are not limited. For example, image data on the Internet may be used as the first training images without being selected. When image data on the Internet is used as the first training images without being selected, images unrelated to the object to be surface inspected may be included. The first training images may be images obtained at at least one of multiple manufacturing sites or multiple production lines where the object is manufactured. The first training images may be, for example, images from the same industry as the object, or images obtained by capturing images of an item made of the same material as the object. Unlike the second training images, the first training images do not require labeling, which will be described in detail later, and therefore can be easily prepared in large quantities. The second training images are images used in the training phase described below for training the downstream task model 13. The second training images are captured images obtained by capturing an image of the surface of an object, and are images that constitute supervised data 220 when combined with labels. The first learning image and the second learning image may be configured from images that are entirely different from each other, or the first learning image and the second captured image may be configured from a portion of the same image. The image acquisition unit 11 is also a functional unit that acquires inspection images, which are captured images obtained by capturing images of objects to be inspected, in the inference phase (the inference phase in which inference is performed using the upstream task model and the downstream task model) described below. That is, when inspecting an object whose surface texture needs to be inspected, such as an object transported to an inspection process, the captured image acquisition unit 11 can also capture an inspection position on the surface of the target object to generate and acquire the captured image. The captured image used for surface inspection acquired in this case is hereinafter referred to as the inspection image. As described below, the inspection image is used to inspect the surface texture of the target object. The upstream task model 12 is a functional unit that includes a base model, which is a deep learning neural network that aims to extract features of an image obtained by capturing an image of the surface of an input object. That is, the upstream task model 12 is configured to include a base model. The base model is a large-scale neural network that has undergone unsupervised learning by inputting an extremely large amount of images and unlabeled data 210 composed of various data, including not only images but also text and sound. The training method for the base model is not particularly limited, and various well-known methods in the field of deep learning can be used. For example, a method that minimizes reconstruction error when a portion of the input image or intermediate feature is masked can be considered. Through such learning, the (trained) base model is expected to be able to extract essential features of the image. Since the upstream task model 12 is equipped with such a base model, when an image is input, it is possible to extract the essential features of the image. Preferably, by configuring the model 12 for the upstream task to learn using the first captured image, which is an image relating to the surface properties of the object at each manufacturing site or manufacturing line, it is also possible to extract with high accuracy the feature quantities relating to the surface properties of the object at each manufacturing site or manufacturing line. Furthermore, the upstream task model 12 is also a functional unit that is trained by inputting first training images in a later-described training phase (a phase in which the upstream task model and the downstream task model are optimized). That is, the upstream task model 12 is configured with a base model, and the base model is trained by unsupervised learning. Therefore, the upstream task model 12 can train the base model using only the first training images, which are unsupervised data 210, without requiring supervised data 220 (which requires manual or other effort to acquire). Therefore, training the upstream task model 12 does not require the effort of acquiring supervised data 220, and therefore the function of extracting image features can be obtained more easily than when using a deep learning model trained by conventional supervised learning. In addition, the upstream task model 12 is also a functional unit that, in the inference phase (a phase in which inference is performed using the upstream task model and the downstream task model) described below, outputs feature quantities corresponding to the test image by inputting the test image. That is, in the inference phase, the upstream task model 12 is composed of a trained base model that has gone through the learning phase and is capable of extracting features, and when inspecting the surface texture of an object, an inspection image of the object to be inspected is input to the trained base model to obtain its features. In other words, the upstream task model 12 can reduce the dimension of the image, which is multidimensional data, to features while maintaining the essential information contained in the inspection image. In other words, the upstream task model 12 receives the inspection image as input and outputs the features extracted from the inspection image. The downstream task model 13 is a functional unit consisting of a deep learning model. That is, the downstream task model 13 also has a deep learning model like the upstream task model 12, but differs in that the upstream task model 12 has a neural network as a base model that outputs features, whereas the downstream task model 13 has a conventional neural network that learns through supervised learning. The downstream task model 13 is also a functional unit that performs supervised learning in a learning phase (a learning phase in which the upstream task model and the downstream task model are optimized) described below by inputting the feature values ​​obtained by inputting the second learning image into the upstream task model 12 and the label indicating the surface condition of the object corresponding to the feature values. The supervised data 220 used for learning by the deep learning model included in the downstream task model 13 is composed of a combination of features obtained from the upstream task model 12 and information (labels) indicating the correct type of surface texture obtained by an operator or inspector checking the surface texture of the object shown in the image used to extract the features. That is, in the learning phase in which the deep learning model is trained, the downstream task model 13 performs inference by inputting the feature obtained by inputting the second learning image into the upstream task model 12 into the deep learning model, and learns by an optimization calculation that minimizes the difference between the inference result and the correct answer (label) corresponding to the feature assigned by an inspector or operator after checking the second captured image. Therefore, after the learning process has progressed, the trained deep learning model provided in the downstream task model 13 will function as an inspection model, capable of inferring the surface condition of an object from the input features. Furthermore, downstream task model 13 is also a functional unit that outputs the surface condition of an object to be inspected by inputting feature quantities corresponding to an inspection image in an inference phase (an inference phase in which inference is performed using the upstream task model and the downstream task model) described below. That is, downstream task model 13 is composed of a deep learning model that can infer the surface condition of an object from the input feature quantities, and when inspecting the surface texture, the feature quantities output by upstream task model 12 based on an inspection image of the object to be inspected are input to the deep learning model (inspection model), thereby inferring and outputting the surface condition of the object. Therefore, the downstream task model 13 executes inference related to surface inspection based on the feature quantities reduced in dimension by the upstream task model 12. That is, the downstream task model 13 receives feature quantities that condense essential information of the image, so the downstream task model 13 can infer the surface condition of the object (especially the type of defect) with high accuracy. As is well known in the field of deep learning, the learning phase refers to the stage of optimizing (learning) the weighting coefficients that make up each layer of a deep learning model that has not yet been optimized. More specifically, in the case of a base model included in upstream task model 12, this refers to the stage where, for example, unsupervised data 210 consisting of a first learning image is input to the base model and optimization operations for the base model are repeated. In the case of an inspection model included in downstream task model 13, this refers to the stage where, for example, supervised data 220 related to the feature values ​​output from upstream task model 12 based on the second learning image is input to the inspection model and optimization operations for the inspection model are repeated. As is well known in the field of deep learning, the inference phase refers to a stage in which inference is made based on input images and feature quantities using a trained deep learning model that has been optimized through the training phase. More specifically, in the case of a base model included in upstream task model 12, it refers to a stage in which, for example, an inspection image is input into the trained base model and feature quantities related to the surface properties of the object depicted in the inspection image are inferred. In the case of an inspection model included in downstream task model 13, it refers to a stage in which, for example, feature quantities output from upstream task model 12 based on the inspection image are input into the trained inspection model and the surface condition of the object depicted in the inspection image is inferred.

[0015] Furthermore, neural network models, including basic models, generally have a multi-stage hierarchical structure and include an extremely large number of weighting coefficients, which makes it difficult to apply them to the high-speed judgments (e.g., real-time judgments) required on manufacturing lines in terms of processing time and hardware load. Furthermore, introducing computational resources to solve these problems increases costs. Therefore, the learning models (upstream task model 12, downstream task model 13) generated in the learning phase may be subjected to model downsizing techniques such as distillation or quantization, which are well known in the field of deep learning, for example, for the entire trained upstream task model 12 and the trained downstream task model 13.

[0016] Distillation is a method for transferring knowledge from a large teacher model to a smaller student model. The teacher model is first trained to generate a probability distribution (soft targets) for each input data. The student model is then trained using this probability distribution, resulting in a lighter, more efficient model that maintains accuracy close to that of the teacher model. Quantization is a technique that converts model parameters and intermediate layer outputs into lower-precision values. Quantization significantly reduces the memory usage and computational complexity of models, allowing them to be used in environments requiring real-time inference or on devices with limited resources.

[0017] As described above, by replacing the trained models (the trained upstream task model 12 (base model) and the trained downstream task model 13 (inspection model)) with miniaturized models using model miniaturization techniques such as distillation or quantization, it is possible to improve the inference speed and the implementability required for surface inspection.

[0018] Furthermore, the surface inspection device 1 may be provided with a surface inspection unit (not shown) in addition to the configuration described above. The surface inspection unit is a functional unit that inspects the object to be inspected based on the surface condition of the object to be inspected output from the downstream task model 13 (i.e., the inspection model). That is, the surface inspection unit judges the pass / fail of the object based on defects that have occurred in the object, such as whether the object can be shipped. The surface inspection unit may also function as an output unit that outputs the estimation results output from the downstream task model 13 in the inference phase. Furthermore, in the above, we have explained an example in which the surface inspection device 1 is a device that performs both the learning phase and the inference phase by itself, but this is not limited to this case, and the surface inspection device 1 may be a device that performs only the learning phase by itself, or the surface inspection device 1 may be a device that performs only the inference phase by itself.

[0019] (Device configuration) The surface inspecting apparatus 1 according to the present disclosure is configured using information devices such as a personal computer. The surface inspecting apparatus 1 may include a processor such as a CPU (Central Processing Unit) and a memory connected to the information devices via a bus, and each function may be ensured by executing a preconfigured control program. The surface inspecting apparatus 1 may also be realized using hardware such as an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), or an FPGA (Field Programmable Gate Array). The control program may also be recorded on a computer-readable recording medium. Examples of computer-readable recording media include portable media such as a flexible disk, a magneto-optical disk, a ROM, and a CD-ROM, and storage devices such as a hard disk built into a computer system. The program may also be transmitted via a telecommunications line.

[0020] (Operation of surface inspection device 1) Next, the operation of the surface inspection apparatus 1 according to this embodiment will be described, divided into a learning phase and an inference phase, with reference to Figures 4A and 4B. The surface inspection apparatus 1 performs the following processes to realize a surface inspection method for inspecting the surface of an object. In the following, an example will be described in which the object is a steel plate manufactured on a production line, and the surface inspection device 1 estimates defects on the surface of the steel plate.

[0021] (Learning phase operation) Fig. 4A is a flowchart showing the operation of the learning phase of the surface inspecting apparatus 1 according to this embodiment. When the surface inspecting apparatus 1 starts the processing of the learning phase shown in Fig. 4A, the process proceeds to step S11.

[0022] In step S11, the image acquisition unit 11 is used to acquire a first learning image and a second learning image, which are captured images obtained by capturing an image of an object (learning image acquisition step).

[0023] The first training image is an image of the surface of a steel plate, which will later become unsupervised data 210 for training a base model. The first training image is an image obtained by imaging an object at a manufacturing site or production line where the object is manufactured. The first training image may be an image of the surface of a steel plate, and the presence or absence of defects and the type of steel plate are not limited. The first training image may be an image including defects other than those determined by the surface inspection device 1, or may include an image of a steel plate of a different type from the steel plate (object) determined by the surface inspection device 1. A steel plate of a different type from the steel plate (object) determined by the surface inspection device 1 is a steel plate that differs from the steel plate (object) determined by the surface inspection device 1 in at least one of chemical composition, metal structure, and manufacturing conditions.

[0024] The second learning image is an image of the surface of the steel plate that is the object, and is an image that will later be combined with a label for learning the inspection model to become supervised data 220. The second learning image is an image that includes a defect that corresponds to one of the defects determined by the surface inspection device 1, and is an image of the steel plate (object) that is to be inspected by the surface inspection device 1. When the processing in step S11 is completed, the process proceeds to step S12.

[0025] In step S12, learning is performed by inputting a first training image into the upstream task model 12 (upstream task model learning step). In step S12, unsupervised learning is performed by inputting the first training image acquired in step S11 into the upstream task model 12 (more specifically, the base model included in the upstream task model 12) as unsupervised data 210. In step S12, unsupervised learning is performed on the upstream task model 12, thereby optimizing the base model included in the upstream task model 12, thereby improving the ability of the upstream task model 12 to reduce the dimension of image data (multidimensional data) into feature data.

[0026] In step S12, images of various steel plate surfaces are input as first learning images to the upstream task model 12. This constructs a trained base model specialized in extracting feature amounts of steel plate surfaces in the upstream task model 12. The construction of the trained upstream task model 12 specialized in extracting feature amounts of steel plate surfaces based on the first learning images is referred to as an upstream task 110 of the learning phase. Here, the base model is trained using the first training image, thereby constructing a trained base model in the upstream task model 12. However, a publicly known base model for image recognition may be obtained from outside and used as the base model that constitutes the upstream task model 12. When the processing in step S12 is completed, the process proceeds to step S13.

[0027] In step S13, the feature amounts obtained by inputting the second learning images into the upstream task model 12 and the labels indicating the surface states of the object corresponding to the feature amounts are input to the downstream task model 13, thereby performing supervised learning (downstream task model learning step). Note that the construction of the trained downstream task model 13 based on the feature amounts obtained by inputting the second learning images into the upstream task model 12 and the labels indicating the surface states of the object corresponding to the feature amounts is referred to as downstream task 120 of the learning phase. That is, in step S13, first, feature amounts obtained in step S12 by inputting second learning images, which are images of the surface of steel plates manufactured on the target production line, into the upstream task model 12 are acquired. Next, an inspector or operator checks the second learning images to acquire labels that are correct answers indicating the surface state of the object depicted in the second learning images (i.e., labels indicating the surface state of the object corresponding to the acquired feature amounts). The acquired feature amounts are then input into the downstream task model 13, whereby an inference result is obtained from the downstream task model 13. An optimization calculation is performed to optimize the weighting coefficients of the downstream task model 13 so that the inference result matches the label corresponding to the feature amounts. By repeating this process, a trained downstream task model 13 suitable for surface inspection of steel plates is constructed. When the process in step S13 ends, the process of the learning phase ends. The processing of steps S11 to S13 described above corresponds to the operation of the learning phase of the surface inspecting apparatus 1. The surface inspecting apparatus 1 may be configured to perform only the operation of the learning phase of steps S11 to S13. Furthermore, after performing the processing of the learning phase of steps S11 to S13, the surface inspecting apparatus 1 may perform the operation of the inference phase of steps S14 to S16 described below following the processing of the learning phase.

[0028] (Inference phase operation) Fig. 4B is a flowchart showing the operation of the inference phase of the surface inspecting apparatus 1 according to this embodiment. When the processing of the inference phase shown in Fig. 4B is started in the surface inspecting apparatus 1, the process proceeds to step S14. In step S14, an inspection image is acquired, which is an image obtained by imaging the object to be inspected, using the image acquisition unit 11 (inspection image acquisition step). The inspection image acquired in step S14 is an image obtained by imaging the surface of the object to be inspected. That is, in step S14, an image to be used for surface inspection is acquired as the inspection image. When the processing in step S14 is completed, the process proceeds to step S15.

[0029] In step S15, an inspection image is input using the upstream task model 12, and feature quantities corresponding to the inspection image are output (upstream task model inference step). That is, in step S15, the inspection image acquired in step S14 is input to the trained upstream task model 12 constructed in step S12 (particularly, the trained base model included in the upstream task model 12), thereby performing inference and outputting feature quantities used to estimate the surface state of the object. The feature quantities output here correspond to information obtained by reducing the dimensionality of the inspection image, which is multidimensional data, while maintaining the essential information contained in the inspection image. When the processing in step S15 is completed, the process proceeds to step S16.

[0030] In step S16, the downstream task model 13 is used to input feature amounts corresponding to the inspection image, and output the surface state of the object to be inspected (downstream task model inference step). That is, in step S16, the feature amounts of the inspection image obtained in the upstream task model inference step S15 are input to the trained downstream task model 13 (particularly, the inspection model included in the downstream task model 13), and the downstream task model 13 estimates the surface state of the object based on the feature amounts. When the processing in step S16 is completed, the processing in the inference phase is completed. The processing of steps S14 to S16 described above corresponds to the operation of the inference phase of the surface inspection apparatus 1. The surface inspection apparatus 1 may perform only the operation of the inference phase of steps S14 to S16, and may perform each inference using the trained upstream task model 12 and the trained downstream task model 13. Furthermore, the surface inspection apparatus 1 may perform the processing of the learning phase of steps S11 to S13, and then perform the operation of the inference phase of steps S14 to S16 following the processing of the learning phase. When the processing of step S16 is completed, the processing of the surface inspection apparatus 1 ends.

[0031] It should be noted that the processing of the surface inspection device 1 does not end when the processing of step S16 is completed, and an object inspection step (not shown) may be performed after the processing of step S16. In the object inspection step, the surface inspection device 1 inspects the object to be inspected using a surface inspection unit based on the surface state of the object output from the downstream task model 13. If the object inspection step is performed, the processing of the surface inspection device 1 ends after the object inspection step is completed.

[0032] (summary) As described above, in the surface inspection apparatus 1 of this embodiment, in the upstream task 110 of the learning phase, the base model included in the upstream task model 12 acquires the ability to reduce the dimension of an image (multidimensional data) to feature quantities while maintaining the essential information contained in the image through unsupervised learning. Furthermore, in the downstream task 120, the inspection model included in the downstream task model 13 is trained using supervised data 220. During this process, the essential information of the multidimensional data contained in the image is not degraded, and only the inspection model in the downstream task 120 requires the supervised data 220, which requires a lot of effort to acquire. Therefore, processing can be performed with a significantly smaller amount of supervised data 220 than conventional deep learning models, thereby achieving surface inspection with the same or higher accuracy as conventional deep learning models, but with less effort than conventional deep learning models.

[0033] Furthermore, conventional deep learning models are specialized learning models for specific manufacturing sites or production lines, making them difficult to transfer to other lines. In contrast, a foundational model trained unsupervised on a variety of images as first training images acquires the ability to extract general-purpose image features. Therefore, even if there are slight differences in images between sites or lines due to differences in background, brightness, image quality, etc., the trained foundational model has acquired the ability to extract general-purpose image features for specific objects, making it possible to extract feature quantities related to the essential characteristics of defects. Therefore, the same trained foundational model can be used at other sites or lines that manufacture similar products.

[0034] Therefore, when the trained model in this embodiment is transferred to another base or line that manufactures similar products, it is only necessary to retrain the inspection model (downstream task model 13) in the downstream task 120 using a small amount of supervised data 220.

[0035] Therefore, according to the deep learning model of this embodiment, the versatility of the base model can be utilized to enable the same trained base model to be used at multiple locations, and at any location, the required inspection performance can be achieved with significantly less supervised data 220 than with conventional deep learning models.

[0036] Furthermore, in the surface inspection device 1 of the present disclosure, captured images obtained by capturing images of objects at multiple manufacturing bases or multiple manufacturing lines may be used as the second learning images. By aggregating data (captured images) from multiple locations and executing downstream task 120, the load required to acquire supervised data 220 can be reduced. For example, in the manufacturing industry, identical or similar products may be manufactured at different locations or on different lines. In such cases, defects to be identified may be common even for products manufactured at different locations or on different lines. Therefore, by aggregating data (captured images) from multiple locations, supervised data 220 can be jointly created at multiple locations or lines. This reduces the amount of supervised data 220 that must be prepared per location compared to preparing all data at a single location.

[0037] Although an embodiment of the present invention has been described above in detail with reference to the drawings, the specific configuration is not limited to this embodiment, and includes designs within the scope of the gist of the present invention. For example, when only the inference phase is performed using a learning model that has already completed learning, the surface inspection device may have a model acquisition unit that acquires a model for an upstream task consisting of a base model that is a deep learning neural network intended to extract features of an image obtained by imaging the surface of the input object, and a model for a downstream task consisting of a deep learning model. [Example]

[0038] Next, examples of the present disclosure will be described, but the conditions in the examples are examples adopted to confirm the feasibility and effects of the present disclosure, and the present disclosure is not limited to these examples. Various conditions may be adopted in the present disclosure as long as they do not deviate from the gist of the present disclosure and the object of the present disclosure is achieved.

[0039] In this example, the downstream task was set to classify six types of typical defects that occur in the steelmaking process. To perform image classification, labels were added to the second learning images in advance. 3,314 images were prepared as evaluation images (inspection images), and these were input to a trained model trained using the following three methods, and the accuracy rate when estimating defects on the steel plate surface was calculated.

[0040] Comparative example: A conventional deep learning model was used, which did not use a base model and was trained entirely on supervised data. 12,412 training images with manually assigned correct labels were input as training data, and a trained model was generated. Invention Example 1: A base model constructed using general domain images was used as the model for the upstream task. For the upstream task, a general-purpose base model (Dino-v2) was used that underwent unsupervised learning using general domain images as first learning images. For the downstream task, the model for the downstream task was additionally trained with supervised data (second learning images) while varying the number of images from 35 to 9,894. Here, the general domain images refer to a wide variety of images available on the Internet, including images unrelated to steel sheets. Invention Example 2: A base model constructed using images from the production line was used as the model for the upstream task. For the upstream task, a base model that had undergone unsupervised learning using approximately 300,000 images taken during the steelmaking process as the first learning images was used. For the downstream task, additional learning was performed on the model for the downstream task using the same method as Invention Example 1.

[0041] FIG. 5 shows a graph illustrating the results of the example. When the learning model of the comparative example was used, 12,412 supervised data were input, and a trained model with an accuracy rate of 85% was generated for 3,314 evaluation images. When using the trained model of Example 1 (a training model including a base model constructed from general images), the number of supervised data required to obtain the same accuracy rate (85%) as the comparative example was approximately 8,000 to 9,000 images. Furthermore, when 9,894 images of supervised data were input in training the downstream task of the training model of Example 1, a trained model with an accuracy rate of 87% was generated. When using the trained model of Example 2 (a training model including a base model constructed from images of the production line), the number of supervised data required to obtain the same accuracy rate (85%) as the comparative example was approximately 350. Furthermore, when 9,894 supervised data were input in training the downstream task of the training model of Example 2, a trained model with an accuracy rate of 92% was generated.

[0042] From these results, it became clear that by constructing a learning model for surface inspection (defect discrimination model) by separating the upstream task of training the base model with unsupervised data and the downstream task of training the inspection model with supervised data, as in the learning model disclosed herein, it is possible to reduce the amount of supervised data required to achieve a predetermined accuracy rate, or to construct a more accurate learning model with the same amount of supervised data as in the past. Furthermore, it became clear that a learning model including a base model constructed with images related to the object (images of the production line in this example), as in Example 2 of the invention, can reduce the amount of supervised data required to achieve a predetermined accuracy rate, or to construct a more accurate learning model, compared to a learning model including a base model constructed with general images. [Explanation of symbols]

[0043] 1. Surface inspection equipment 11 Image acquisition unit 12 Model for upstream tasks 13 Models for downstream tasks 110 Upstream Tasks 120 Downstream Tasks 210 Unsupervised Data 220 Supervised Data 400 Deep Learning Models (Conventional)

Claims

1. A surface inspection device for inspecting a surface of an object, comprising: an image acquisition unit that acquires a captured image obtained by capturing an image of the surface of the object; an upstream task model consisting of a base model that is a deep learning neural network intended to extract features of an image obtained by imaging the surface of the input object; a model for downstream tasks consisting of a deep learning model; and In a learning phase in which the upstream task model and the downstream task model are optimized, the image acquisition unit acquires a first learning image and a second learning image, which are captured images obtained by capturing images of the object; the upstream task model is trained by inputting the first training image; the downstream task model is trained by inputting feature amounts obtained by inputting the second training image into the upstream task model and labels indicating the surface state of the object corresponding to the feature amounts; In an inference phase in which inference is performed using the upstream task model and the downstream task model, the image acquisition unit acquires an inspection image, which is an image obtained by capturing an image of the object to be inspected; the upstream task model receives the test image as input and outputs a feature quantity corresponding to the test image; a downstream task model that outputs a surface state of the object to be inspected by inputting the feature amount corresponding to the inspection image;

2. A surface inspection device for inspecting a surface of an object, comprising: an image acquisition unit that acquires a captured image obtained by capturing an image of the surface of the object; an upstream task model consisting of a base model that is a deep learning neural network intended to extract features of an image obtained by imaging the surface of the input object; a model for downstream tasks consisting of a deep learning model; and In a learning phase in which the upstream task model and the downstream task model are optimized, the image acquisition unit acquires a first learning image and a second learning image, which are captured images obtained by capturing images of the object; the upstream task model is trained by inputting the first training image; the downstream task model is trained by inputting feature amounts obtained by inputting the second training image into the upstream task model, and labels indicating the surface condition of the object corresponding to the feature amounts.

3. A surface inspection device for inspecting a surface of an object, comprising: an image acquisition unit that acquires a captured image obtained by capturing an image of the surface of the object; a model acquisition unit that acquires an upstream task model consisting of a base model that is a deep learning neural network intended to extract features of an image obtained by imaging the surface of the input object, and a downstream task model consisting of a deep learning model; and The model acquisition unit acquiring a first learning image and a second learning image, which are captured images obtained by capturing an image of the object, and acquiring the upstream task model trained by inputting the first learning image; and acquiring the downstream task model trained by inputting feature amounts obtained by inputting the second learning image into the upstream task model and labels indicating the surface state of the object corresponding to the feature amounts; In an inference phase in which inference is performed using the upstream task model and the downstream task model, the image acquisition unit acquires an inspection image, which is an image obtained by capturing an image of the object to be inspected; the upstream task model receives the test image as input and outputs a feature quantity corresponding to the test image; a downstream task model that outputs a surface state of the object to be inspected by inputting the feature amount corresponding to the inspection image;

4. 4. The surface inspection device according to claim 1, wherein the upstream task model and the downstream task model are replaced by models that are miniaturized by using a distillation or quantization technique in a deep learning model.

5. A surface inspection method for inspecting a surface of an object, comprising: an image acquisition unit that acquires a captured image obtained by capturing an image of the surface of the object; an upstream task model consisting of a base model that is a deep learning neural network intended to extract features of an image obtained by imaging the surface of the input object; a model for downstream tasks consisting of a deep learning model; Using a surface inspection device having In a learning phase in which the upstream task model and the downstream task model are optimized, a learning image acquisition step of acquiring, using the image acquisition unit, a first learning image and a second learning image, which are captured images obtained by capturing images of the object; an upstream task model learning step in which the first learning image is input to the upstream task model to perform learning; a downstream task model learning step in which the downstream task model is learned by supervised learning by inputting, into the downstream task model, feature amounts obtained by inputting the second learning image into the upstream task model and labels indicating the surface state of the object corresponding to the feature amounts; In an inference phase in which inference is performed using the upstream task model and the downstream task model, an inspection image acquisition step of acquiring an inspection image, which is an image obtained by imaging the object to be inspected using the image acquisition unit; an upstream task model inference step of inputting the test image using the upstream task model and outputting feature quantities corresponding to the test image; a downstream task model inference step of outputting a surface state of the object to be inspected by inputting the feature amount corresponding to the inspection image using the downstream task model; A surface inspection method comprising:

Citation Information

Patent Citations

  • Information processing apparatus, information processing method, and program

    JP2018120300A

  • Vehicle damage estimation device, estimation program therefor and estimation method therefor

    JP2021167991A

  • Anomaly detection system, learning apparatus, anomaly detection program, learning program, anomaly detection method, and learning method

    JP2022015575A

  • Foreign matter detection device, foreign matter detection system, and foreign matter detection method

    JP2022066637A

  • System and method for applying deep learning tools to machine vision and interface therefor

    JP2024541040A