Fault detection method and apparatus, model training method and apparatus, and electronic device
By automatically configuring and adjusting the training parameters of the deep learning algorithm model, the problem of manual intervention in the existing technology is solved, and the accuracy and efficiency of defect detection are improved.
Patent Information
- Application Number
- JP2022571810
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-01-28
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2041-01-28
AI Technical Summary
The prior art requires manual intervention in parameter adjustment of deep learning algorithm models, resulting in waste of resources and human errors, affecting detection accuracy.
By obtaining the sample data set, identifying its characteristic information, configuring training parameters, using the sample data set to train the initial model to obtain the target model, and then use it to detect defects in the actual product.
It reduces manual intervention, improves the accuracy and efficiency of defect detection, and avoids waste of resources caused by human errors.
Smart Images

Figure 0007678826000003 
Figure 0007678826000004 
Figure 0007678826000005
Abstract
Description
[Technical field]
[0001] The present disclosure relates to the technical field of defect detection, and in particular to a defect detection method and apparatus thereof, a model training method and apparatus thereof, a computer-readable recording medium, and an electronic device. [Background technology]
[0002] In the field of screen production, defects occur in the produced products due to problems in the process such as devices, parameters, operations, and environmental interference. With the rise of artificial intelligence algorithms, such as deep learning, the use of deep algorithm learning models to detect defects is becoming more widespread.
[0003] However, in the prior art, manual tuning is often required for parameter adjustment in deep algorithm learning models, which may waste human resources and cause losses due to human malfunction.
[0004] It should be noted that the information disclosed in the above Background section is intended solely to facilitate understanding of the context of the present disclosure, and may therefore include information that does not constitute prior art known to those skilled in the art. Summary of the Invention [Problem to be solved by the invention]
[0005] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by the practice of the present disclosure. [Means for solving the problem]
[0006] A first aspect of the present disclosure provides a defect detection method, comprising: Obtaining a sample data set including defective product data and identifying characteristic information of the sample data set; Obtaining the initial model, configuring training parameters based on the characteristic information; training the initial model using the sample data set based on the training parameters to obtain a target model; inputting actual data of a product corresponding to the sample data set into the target model to obtain defect information of the product; the feature information includes a number of samples in the sample data set; the initial model is a neural network model; The training parameters include at least one of a learning rate descent policy, a total number of training rounds, and a testing policy, the learning rate descent policy includes a number of times the learning rate is descent and a number of rounds during the descent, and the testing policy includes a number of times testing is performed and a number of rounds during testing.
[0007] A second aspect of the present disclosure provides a model training method, comprising: Obtaining a sample data set including defective product data and identifying characteristic information of the sample data set; Obtaining the initial model, configuring training parameters based on the characteristic information; training the initial model based on the training parameters using the sample data set to obtain a target model; the feature information includes a number of samples in the sample data set; the initial model is a neural network model; the target model is used to perform defect detection on actual data of a product corresponding to the sample data set; The training parameters include at least one of a learning rate descent policy, a total number of training rounds, and a testing policy, the learning rate descent policy includes a number of times the learning rate is descent and a number of rounds during the descent, and the testing policy includes a number of times testing is performed and a number of rounds during testing.
[0008] A third aspect of the present disclosure provides a model training method, comprising: acquiring a sample data set including defective product data in response to a configuration operation by a user on parameters of the sample data set; and identifying characteristic information of the sample data set; Obtaining the initial model, configuring training parameters based on the characteristic information and generating a display interface for the training parameters; training the initial model based on the training parameters using the sample data set to obtain a target model; the feature information includes a number of samples in the sample data set; the initial model is a neural network model; the target model is used to perform defect detection on actual data of a product corresponding to the sample data set; The training parameters displayed in the training parameter display interface include at least one of a learning rate descent policy, a total number of training rounds, and a test policy, the learning rate descent policy includes the number of times the learning rate is descent and the number of rounds during descent, and the test policy includes the number of tests and the number of rounds during testing.
[0009] A fourth aspect of the present disclosure provides a detection system including a data management module, a training management module, and a model management module, wherein: the data management module is configured to store and manage the sample data; a training management module configured to perform any of the defect detection methods described above, any of the model training methods described above, or any of the model training methods described above; A model management module is configured to store, display and manage the target models.
[0010] A fifth aspect of the present disclosure provides a computer-readable recording medium storing a computer program, wherein the computer program, when executed by a processor, realizes any of the defect detection methods described above, any of the model training methods described above, or any of the model training methods described above.
[0011] A sixth aspect of the present disclosure provides an electronic device including a processor and a memory storing one or more programs, wherein: The one or more programs, when executed by the one or more processors, cause the one or more processors to perform any of the defect detection methods described above, any of the model training methods described above, or any of the model training methods described above.
[0012] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure.
[0013] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure. Apparently, the drawings in the following description are only some embodiments of the present disclosure, and those skilled in the art can also obtain other drawings from these drawings without creative efforts. [Brief description of the drawings]
[0014] [Figure 1] FIG. 1 is a schematic diagram of a system configuration to which an embodiment of the present disclosure is applied. [Diagram 2] FIG. 1 is a schematic diagram of an electronic device to which an embodiment of the present disclosure can be applied. [Diagram 3] 1 is a flowchart of a defect detection method according to an embodiment of the present disclosure. [Figure 4] FIG. 2 is a schematic diagram of a loss curve in an embodiment of the present disclosure. [Diagram 5] 1 is a flowchart of a model training method in an embodiment of the present disclosure. [Figure 6] 13 is a flowchart of another model detection method according to an embodiment of the present disclosure. [Figure 7] FIG. 13 is a schematic diagram of a display interface of training parameters in an embodiment of the present disclosure. [Figure 8] FIG. 13 is a selection interface diagram for deciding whether to train a defect in an embodiment of the present disclosure. [Figure 9] FIG. 2 is a schematic diagram of a training process in an embodiment of the present disclosure. [Figure 10] FIG. 1 is a schematic diagram of a confusion matrix in an embodiment of the present disclosure. [Figure 11] FIG. 13 is an interface diagram of a target model in an embodiment of the present disclosure. [Figure 12] FIG. 1 is a schematic diagram of a configuration of a defect detection system according to an embodiment of the present disclosure. [Figure 13] FIG. 1 is a configuration diagram of a defect detection system according to an embodiment of the present disclosure. [Figure 14] FIG. 13 is a schematic diagram of a data set preparation interface in an embodiment of the present disclosure. [Figure 15] FIG. 13 is a diagram illustrating a management interface for a training dataset in an embodiment of the present disclosure. [Figure 16] FIG. 13 is a schematic diagram of an interface for establishing a data set in an embodiment of the present disclosure. [Figure 17] FIG. 13 is a diagram illustrating an interface for displaying detailed information about a data set in an embodiment of the present disclosure. [Figure 18] FIG. 13 is an interface diagram of model management in an embodiment of the present disclosure. [Figure 19] FIG. 13 is an interface diagram for establishing a training task in an embodiment of the present disclosure. [Figure 20] FIG. 13 is an interface diagram for changing training parameters in an embodiment of the present disclosure. [Figure 21] FIG. 13 is a schematic diagram of a display interface of a model training process in an embodiment of the present disclosure. [Figure 22]FIG. 2 is a schematic diagram of data direction of a defect detection system in accordance with an embodiment of the present disclosure. [Figure 23] FIG. 1 is a schematic diagram of a configuration of a defect detection device according to an embodiment of the present disclosure. [Figure 24] FIG. 1 is a schematic diagram of a configuration of a model training device in an embodiment of the present disclosure. [Diagram 25] FIG. 13 is a schematic diagram of a configuration of another model training device in an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0015] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, exemplary embodiments may be embodied in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the concept of exemplary embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0016] In addition, the accompanying drawings are merely schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same notations in the drawings indicate the same or similar parts, and therefore their repeated description is omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically separated entities. These functional entities may be implemented in the form of software, in one or more hardware modules or integrated circuits, or in different network devices and / or processor devices and / or microcontroller devices.
[0017] FIG. 1 is a schematic diagram of a system configuration of an application environment of a defect detection method and apparatus to which an embodiment of the present disclosure is applied.
[0018] As shown in FIG. 1, the system configuration 100 may include one or more of terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, among others. The terminal devices 101, 102, 103 may be various electronic devices having defect detection capabilities, such as, but not limited to, desktop computers, portable computers, smartphones, tablet terminals, and the like. Note that the number of terminal devices, networks, and servers in FIG. 1 is merely a schematic diagram. There may be any number of terminal devices, networks, and servers according to the needs of the implementation. For example, the server 105 may be a server cluster, etc., consisting of multiple servers.
[0019] The defect detection method in the embodiment of the present disclosure is executed by the terminal device 101, 102, 103, and correspondingly, the defect detection device is provided in the terminal device 101, 102, 103. Those skilled in the art will understand that the defect detection method in the embodiment of the present disclosure may also be executed by the server 105, and correspondingly, the defect detection device is provided in the server 105, and is not limited to this embodiment. For example, in one embodiment, a user obtains a sample dataset including defective product data through the terminal device 101, 102, 103, identifies feature information of the sample dataset, and the feature information includes the number of samples of the sample dataset, and then uploads the raw depth data sample dataset to the server 105, and the server obtains an initial model, configures training parameters according to the feature information, uses the sample dataset to train the initial model according to the training parameters to obtain a target model, inputs the actual data of the defective product corresponding to the sample dataset into the target model to obtain defect information of the defective product, and transmits the defect information to the terminal device 101, 102, 103, etc.
[0020] An exemplary embodiment of the present disclosure provides an electronic device for performing the defect detection method, which may be the terminal device 101, 102, 103 or the server 105 of Fig. 1. The electronic device includes at least a processor and a memory for storing executable instructions of the processor, and the processor performs the defect detection method by executing the executable instructions.
[0021] In one embodiment of the present disclosure, the system configuration may be a distributed product defect analysis system, which may be a group of computers connected together via a network to transmit and communicate messages and then coordinate their operations to form a system. The components interact with each other to achieve a common goal. The network may be an Internet of Things (IoT) based Internet and / or communication network, which may be wired or wireless, such as a Local Area Network (LAN), a Metropolitan Area Network (MAN), a Wide Area Network (WAN), a cellular data communication network, or any other electronic network capable of exchanging information. A distributed computing system may have software components, such as software objects, or other types of individually addressable isolated entities, such as distributed objects, agents, actors, virtual components, etc. Typically, such components are individually addressable and have a unique ID (such as an integer, GUID, string, opaque data structure, etc.) within the distributed system. In a distributed system that allows geographical distribution, applications may reside in clusters. There are various systems, components, and network configurations that support distributed computing environments. For example, computing systems may be connected by wired or wireless systems, by local networks or widely distributed networks. Currently, many networks are coupled to the Internet, which provides an infrastructure for widely distributed computing, and includes many different networks, but any network infrastructure may be used for communications liable to occur in the systems described in the various embodiments, for example.
[0022] A distributed product defect analysis system provides for the sharing of computer resources and services through communication exchange between computing devices and systems. These resources and services may include the exchange of objects such as files, cache storage, disk storage, and the like. These resources and services may also include sharing of processing power across multiple processing units for load balancing, resource scaling, processing specialization, and the like. For example, a distributed product defect analysis system may include hosts with network topologies and network infrastructures such as client device / server, peer-to-peer, or hybrid architectures.
[0023] The configuration of the electronic device will be described below by taking the mobile terminal 200 of FIG. 2 as an example. It should be understood by those skilled in the art that the structure of FIG. 2 can be applied to fixed-type equipment as well as parts used specifically for mobile purposes. In other embodiments, the mobile terminal 200 may be configured with more or fewer components than those shown, or may be configured with a combination of certain components, or a division of certain components, or a different arrangement of components. The components shown in the figures may be implemented as hardware, software, or a combination of software and hardware. The interface connection relationships between the components are merely schematic and do not constitute structural limitations of the mobile terminal 200. In other embodiments, the mobile terminal 200 may be implemented with interface connections or combinations of interface connections different from those shown in FIG. 2.
[0024] 2, the mobile terminal 200 may specifically include a processor 210, an internal memory 221, an external memory interface 222, a Universal Serial Bus (USB) interface 230, a charging management module 240, a power management module 241, a battery 242, an antenna 1, an antenna 2, a mobile communication module 250, a wireless communication module 260, a voice module 270, a speaker 271, a receiver 272, a microphone 273, a headphone interface 274, a sensor module 280, a display 290, a camera module 291, an indicator 292, a motor 293, a key 294, and a subscriber identification module (SIM) card interface. Here, the sensor module 280 may be composed of a depth sensor 2801, a pressure sensor 2802, a gyro sensor 2803, etc.
[0025] The processor 210 may include one or more processing units, for example, the processor 210 may include an Application Processor (AP), a modem processor, a Graphics Processing Unit (GPU), an Image Signal Processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a Neural-Network Processing Unit (NPU), where different processing units may be independent devices or may be integrated into one or more processors.
[0026] The NPU is a neural network (NN) computing processor that can process input information at high speed and continuously self-learn by referring to the structure of biological neural networks, such as the communication patterns between neurons in the human brain. The NPU enables intelligent cognitive applications such as image recognition, face recognition, voice recognition, and text understanding in the mobile terminal 200.
[0027] The processor 210 is provided with a memory. The memory can store instructions for implementing six module functions, including a detection instruction, a connection instruction, an information management instruction, an analysis instruction, a data transfer instruction, and a notification instruction, and the instructions are executed and controlled by the processor 210.
[0028] The charging management module 240 is used to receive charging input from a charger. The power management module 241 is used to connect the battery 242, the charging management module 240, and the processor 210. The power management module 241 receives input from the battery 242 and / or the charging management module 240, and supplies power to the processor 210, the internal memory 221, the display 290, the camera module 291, the wireless communication module 260, etc.
[0029] The wireless communication function of the mobile terminal 200 can be realized by an antenna 1, an antenna 2, a mobile communication module 250, a wireless communication module 260, a modem processor, and a baseband processor. Here, the antenna 1 and the antenna 2 are used to transmit and receive electromagnetic wave signals, the mobile communication module 250 may provide solutions for wireless communication including 2G / 3G / 4G / 5G etc. applied to the mobile terminal 200, the modem processor may include a modulator and a demodulator, and the wireless communication module 260 provides wireless communication solutions such as WLAN (Wireless Local Area Networks) (e.g., Wi-Fi (Wireless Fidelity) network), Bluetooth (BT), and others applied to the mobile terminal 200. In some embodiments, the antenna 1 of the mobile terminal 200 is coupled to the mobile communication module 250, and the antenna 2 is coupled to the wireless communication module 260, allowing the mobile terminal 200 to communicate with networks as well as other devices through wireless communication technology.
[0030] The mobile terminal 200 realizes display functions through a GPU, a display 290, an application processor, etc. The GPU is a microprocessor for image processing that connects the display 290 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 210 may be composed of one or more GPUs that execute program instructions that generate or modify display information.
[0031] The mobile terminal 200 can realize a photographing function through an ISP, a camera module 291, a video codec, a GPU, a display 290, and an application processor. Here, the ISP is used to process data fed back from the camera module 291, which is used to take still or moving images, the digital signal processor is used to process digital signals and can process other digital signals in addition to digital image signals, and the video codec is used to compress or decompress digital videos, and the mobile terminal 200 can also support one or more video codecs.
[0032] By connecting an external memory card such as a MicroSD card using the external memory interface 222, the storage capacity of the mobile terminal 200 can be expanded. The external memory card communicates with the processor 210 via the external memory interface 222 to realize a data storage function. For example, files such as music and videos are stored on the external memory card.
[0033] The internal memory 221 may be used to store computer executable program code, which includes instructions. The internal memory 221 may be configured with a program storage area and a data storage area. Here, the program storage area may store an operating system, an application required for at least one function (e.g., a voice playback function, an image playback function, etc.), etc. The data storage area may store data (e.g., voice data, a phone book, etc.) created during use of the mobile terminal 200. Furthermore, the internal memory 221 may include a high-speed random access memory, and may also include at least one non-volatile memory such as a disk memory device, a flash memory device, or a Universal Flash Storage (UFS). The processor 210 executes the instructions stored in the internal memory 221 and / or the instructions stored in a memory provided in the processor to execute various function applications of the mobile terminal 200 and also perform data processing.
[0034] The mobile terminal 200 can implement audio functions, such as music playback, recording, etc., via an audio module 270, a speaker 271, a handset 272, a microphone 273, a headphone interface 274, and an application processor.
[0035] The depth sensor 2801 is used to obtain depth information of a scene. In some embodiments, the depth sensor may be provided in the camera module 291.
[0036] The pressure sensor 2802 can be used to sense a pressure signal and convert the pressure signal into an electrical signal. In some embodiments, the pressure sensor 2802 can be provided in the display 290. There are many types of pressure sensor 2802, such as resistive pressure sensor, inductive pressure sensor, capacitive pressure sensor, etc.
[0037] The gyro sensor 2803 may be used to determine the motion attitude of the mobile terminal 200. In some embodiments, the angular velocity of the mobile terminal 200 about three axes (i.e., x, y, and z axes) may be determined by the gyro sensor 2803. The gyro sensor 2803 may be used for anti-shake, navigation, photography, physical gaming scenarios, and the like.
[0038] In addition, according to actual needs, other functional sensors can be provided in the sensor module 280, such as an air pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a proximity light sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, a bone conduction sensor, etc.
[0039] Other devices providing auxiliary functions may also be included in the mobile terminal 200. For example, the keys 294 may include a power on key, volume keys, etc., and a user may use key inputs to generate key signal inputs related to user settings and function control of the mobile terminal 200. Further examples include an indicator 292, a motor 293, a SIM card interface 295, etc.
[0040] In the field of screen manufacturing, in terms of related technology, there are problems in the processes such as equipment, parameters, operation, and environmental interference, which results in defects in the produced products. After each process, a large amount of image data is generated after optical inspection (AOI), and specialized operators need to judge the level of these images as defective products. With the rise of artificial intelligence algorithms, which are representative of deep learning, AI algorithms are introduced into the process of judging the level of defective images, creating an automatic defective image detection system (ADC).
[0041] The ADC system is composed of four subsystems: data mark system, GPU server (algorithm system), TMS system, and training system. In order to simplify business processes and save hardware resources, the first three of these subsystems are combined to realize the function of automatic detection of defective images on the production line by AI algorithm inference service, that is, the above systems can work normally even without the training system. However, these systems cannot update the algorithm model, and if an AI model needs to be updated, it needs to be developed and deployed by the algorithm developer. The role of the training system is to facilitate the training of algorithms during project development and to facilitate the updating of models during project operation and maintenance periods.
[0042] In the factory, the production process and AOI devices are constantly adjusted, and the deep learning algorithm is a data-driven technology, so the AOI image changes due to constant adjustments of the production process and devices, which reduces the accuracy of the algorithm model. Meanwhile, when the production of a new product starts, the model also needs to be re-adjusted to accommodate the different AOI images corresponding to the new product.
[0043] To improve the robustness of the ADC system, the system can train a new model even after the AOI image changes, ensuring the accuracy of the algorithm. Models trained by deep learning algorithms usually consist of a large number of training parameters, which often need to be manually adjusted for different images. This can waste human resources and cause losses due to human malfunctions.
[0044] In view of the above shortcomings, the present disclosure first provides a defect detection method, as shown in FIG. 3, the defect detection method includes the following steps: In step S310, a sample data set including defective product data is obtained, and characteristic information of the sample data set is identified, the characteristic information including a number of samples of the sample data set; In step S320, an initial model is obtained, and the initial model is a neural network model; In step S330, training parameters are configured based on the characteristic information; In step S340, the sample data set is used to train the initial model based on the training parameters to obtain a target model; In step S350, actual data of the product corresponding to the sample data set is input into the target model to obtain defect information of the product.
[0045] Here, the training parameters include at least one of a learning rate descent policy, a total number of training rounds, and a test policy, the learning rate descent policy includes the number of times the learning rate is descent and the number of rounds during descent, and the test policy includes the number of times testing is performed and the number of rounds during testing.
[0046] Compared with the related art, the technical solution in this embodiment determines the training parameters of the model according to the feature information obtained from the training data, and uses the number of samples in the feature information to determine the learning rate descent policy, the total number of training rounds and the test policy in the training parameters, and does not involve any human operation, so that it can save human resources and avoid losses caused by human malfunction. Meanwhile, the learning rate descent policy includes the number of learning rate descents and the number of rounds during descent, the test policy includes the number of tests and the number of rounds during testing, and configures the number of learning rate descents and the number of rounds during descent, the number of tests and the number of rounds during testing, and the learning rate descent policy and the test policy have a great impact on corresponding to defect detection, so by configuring the above training parameters, the accuracy of the obtained target model for defect detection can be greatly improved.
[0047] In step S310, a sample data set including defective product data is obtained, and characteristic information of the sample data set is identified, where the characteristic information includes a number of samples of the sample data set.
[0048] In this embodiment, first, a sample dataset is obtained, and characteristic information of the sample dataset is identified, specifically, the number of samples in the sample dataset is obtained, and the type, size, etc. of pictures of defective products included in the sample dataset are obtained, and this embodiment is not limited thereto, where the sample dataset may include defective product data, and where the defective product data may include pictures of defective products or other product data of defective products, but this embodiment is not limited thereto.
[0049] In this embodiment, here, the feature information of the above sample dataset may include the type, size, etc. of the picture of the defective product in the sample data, and may include the number of samples of the above sample data, where the number of samples may be 10,000, 20,000, etc., and may be defined according to the user's requirements, but is not limited to this in this embodiment.
[0050] In step S320, an initial model is obtained, and the initial model is a neural network model; In this embodiment, the initial model may be, but is not limited to, a convolutional neural network (CNN) model, a target detection convolutional neural network (faster-RCNN) model, a recurrent neural network (RNN) model, a generative adversarial network (GAN) model, or other neural network models known to those skilled in the art.
[0051] In this embodiment, the above-mentioned initial model can be determined according to the type of defect picture in the above-mentioned defective product.Specifically, in this embodiment, different or the same initial model can be selected according to the image of process or product type according to actual service requirements, for example, when the image in the sample data set is the image of intermediate site, the initial model is a convolutional neural network (CNN) model, when the image in the sample data set is the image of final site, the initial model is a convolutional neural network (CNN) model, target detection convolutional neural network (faster-RCNN) model, which is not particularly limited in this embodiment.
[0052] In step S330, training parameters are configured based on the characteristic information; In this embodiment, the training parameters include a learning rate descent policy, a total number of training rounds, and a test policy, where configuring the training parameters based on the above feature information may be configuring a learning rate descent policy, a total number of training rounds, and a test policy based on the number of samples in the feature information, where the learning rate descent policy includes the number of learning rate descents and the number of rounds during descent, and the test policy includes the number of tests and the number of rounds during testing.
[0053] Specifically, the total number of rounds of the training is positively correlated with the number of samples. For example, when the number of samples is less than or equal to 10000, the total number of rounds of the training is configured as 300000. When the number of samples is greater than 10000, the total number of rounds of the training is configured using the following formula for determining the total number of rounds, where the formula for determining the total number of rounds is Y=300000+INT(X / 10000)×b, where Y indicates the total number of rounds of the training, X indicates the number of samples, X is greater than or equal to 10000, INT is a rounding function, b is a growth factor and a fixed value, and b is greater than or equal to 30000 and less than or equal to 70000. In this embodiment, the value of b may be 50000 or 60000, but is not limited to this in this embodiment. In this embodiment, the mapping relationship between the number of samples and the total number of rounds of the training may be the result of optimization obtained by multiple tests, or may be defined according to the user's request, but is not limited to this in this embodiment.
[0054] In this embodiment, the number of rounds during the learning rate drop is positively correlated with the total number of training rounds, where the number of rounds during testing is equal to or greater than the number of rounds during the first learning rate drop and is equal to or less than the total number of training rounds, the number of times the learning rate is dropped is multiple, and at least two tests are performed within a predetermined range of the number of rounds of the second learning rate drop, and two tests may be performed, or three or more tests may be performed, but this embodiment is not limited thereto. Multiple learning rate drops are performed during training, and the number of drops with the best results after multiple drops is selected to improve the accuracy of the obtained target model, and further improve the accuracy of defect detection, and multiple tests are performed on the model during the training process, and the model with the best test result is selected as the target model, and further improve the accuracy of defect detection.
[0055] In this embodiment, the above learning rate drop method may be a division constant drop, an exponential drop, a natural exponential drop, a cosine drop, etc., but is not limited thereto in this embodiment; here, the above learning rate drop width is related to the above learning rate drop and is related to each parameter in the configured above learning rate drop method, and may be directly set to, for example, 0.1, 0.05, etc., but is not particularly limited thereto in this embodiment.
[0056] In one embodiment of the present disclosure, the above feature information includes a size and type of pictures of the defective product in the sample dataset, and configuring training parameters based on the above feature information further includes configuring a size of an input image to be input to the initial model based on the size and type of the pictures.
[0057] Specifically, when the picture type of the defective product is an AOI color image or a DM image, the size of the input image is a first preset multiple of the picture size of the defective product, and when the picture type of the defective product is a TDI image, the size of the input image is a second preset multiple of the picture size of the defective product, the first preset multiple is less than or equal to 1, and the second preset multiple is greater than or equal to 1.
[0058] In this embodiment, the above-mentioned first preset multiple may be 0.25 or more and 0.6 or less, and the second preset multiple may be 3 or more and 6 or less, for example, the technical specification of the size of the input image is mainly determined by the name of the data set, that is, the image type and site to which the data set belongs, etc., where, for the AOI color image of the SD&Final&mask site, the original picture size is 2000*2000 on average, so the size of the input picture is either 500, 688, 864, 1000, or 1200. For the grayscale image of tdi, the original picture size is 64*64 on average, so the size of the input picture is either 192, 208, 224, 240, or 256. The size of the input picture can also be customized according to the needs of the user, and is not particularly limited in this embodiment.
[0059] In another embodiment, for example, the technical specification of the size of the input picture is mainly determined by the name of the dataset, i.e., the image type and site to which the dataset belongs, etc., where the AOI color image of the SD&Final&mask site has an original picture size of 2000*2000 on average, so the input picture size is a plurality of 500, 688, 864, 1000, 1200. The grayscale image of tdi has an original picture size of 64*64 on average, so the input picture size is a plurality of 192, 208, 224, 240, 256. The number of input images is greater than the number of original images, and the number of samples may be the number of the above input images.
[0060] In one embodiment of the present disclosure, the feature information further includes a defect level of the above-mentioned defective product and a number of samples corresponding to each type of defect, and the training parameters further include a confidence level, where the confidence level in the training process can be configured based on the number of samples corresponding to each type of defect and the defect level.
[0061] Specifically, first, a predetermined number is set, and the number of samples corresponding to the above various defects and the magnitude of the predetermined number are determined, and if the number of samples corresponding to the above defects is greater than the predetermined number, a reliability is configured based on the defect level, where the defect level includes a first defect level and a second defect level, and if the defect level of the above defects is the first defect level, the reliability is configured as a first reliability, and if the defect level of the above defects is the second defect level, the reliability is configured as a second reliability, where the second reliability is greater than the first reliability.
[0062] In this embodiment, the above-mentioned predetermined number may be 50, 100, etc., and may be defined according to the user's request, but is not limited to this in this embodiment. Here, the first reliability is 0.6 or more and 0.7 or less, and the second reliability is 0.8 or more and 0.9 or less. The above-mentioned first reliability sum and second reliability value are defined according to the user's request, but are not limited to this in this embodiment.
[0063] For example, for defects with a high occurrence rate and low importance, i.e., defects with a low defect level, a low reliability can be set, for example, for PI820 with no defect and PI800 with a minor defect, the reliability is set to 0.6, i.e., in this figure, it can be determined that there is a defect if the probability score for PI800 or PI820 exceeds 0.6. For defects with a low occurrence rate and high importance, i.e., defects with a high defect level, a higher reliability can be set, for example, for GT011 and SD011, which are serious defects, the reliability is set to 0.85, i.e., in this figure, it is determined that there is a defect only if the probability score for GT011 or SD011 exceeds 0.6. For other figures with low reliability, all are determined to be unknow (not recognized by AI), and manual processing is performed to prevent missed determinations.
[0064] In one embodiment of the present disclosure, after configuring the above training parameters based on the above feature information, the above method further includes generating a display interface of one training parameter, and a parameter modification icon is provided in the display interface of the training parameter, and after a user triggers the parameter modification icon, modifiable parameters are displayed, and the user can modify the above configured training parameters in the modification interface.
[0065] In step S340, the sample data set is used to train the initial model based on the training parameters to obtain a target model; In one embodiment of the present disclosure, after configuration and modification to the above training parameters, the above obtained initial model can be trained using the above sample dataset to obtain a target model.
[0066] The target model is a neural network model that is primarily based on deep learning. For example, the target model may be based on a feedforward neural network. A feedforward network can be implemented as a ringless diagram with nodes arranged in layers. Typically, a feedforward network topology includes an input layer and an output layer separated by at least one hidden layer. The hidden layer converts the input received by the input layer into a representation that helps generate the output at the output layer. The network nodes are fully connected to the nodes of adjacent layers through edges, but there are no edges between the nodes within each layer. Data received at the nodes of the input layer of the feedforward network are propagated (i.e., "feed forward") to the nodes of the output layer through an activation function that calculates the state of the nodes of each successive layer of the network based on coefficients ("weights"), respectively associated with each edge connecting these layers. The output of the target model can take various forms that are not limited by this disclosure. The target model may also include other neural network models, such as, but not limited to, a convolutional neural network (CNN) model, a recurrent neural network (RNN) model, a generative adversarial network (GAN) model, or other neural network models known to those skilled in the art.
[0067] The training with sample data may include the steps of selecting a network topology, using a training data set representative of the problem to be modeled by the network, and adjusting the weights until the network model operates with minimal error for all instances of the training data set. For example, in the training process of supervised learning of a neural network, the outputs generated by the network in response to inputs representing instances of the training data set are compared with outputs marking the instances as "correct", an error signal representing the difference between the outputs and the marked outputs is calculated, and the weights associated with the connections are adjusted to minimize the error as the error signal is propagated backwards through the layers of the network. When the error for each output generated from an instance of the training data set is minimized, the initial model is considered "trained" and is defined as the target model.
[0068] In this embodiment, during the training of the initial model, the defect detection method further includes obtaining a loss curve during the training process, and then the training parameters may be adjusted according to the loss curve. Specifically, referring to FIG. 4, the abscissa of the loss curve is the number of training rounds, and the ordinate is the loss value. During the model training process, the loss curve is updated in real time according to the state during training, and the user can observe the loss curve and adjust the training parameters according to the state. Specifically, if the loss curve is always disturbed and does not show a downward trend, it is determined that the training parameters are not properly configured, and the training should be stopped, and the learning rate parameter and the learning rate descent policy should be readjusted and retrained. If the loss curve gradually drops, it should be continuously observed and the training should be stopped or the initial learning rate should be increased in the next training. If the loss curve still shows a downward trend after the training is completed (normally it should be smoothed out eventually), the maximum number of training rounds should be increased and retraining should be performed after the training is completed.
[0069] In one embodiment of the present disclosure, according to the test policy, a model that has reached the number of training rounds and the number of rounds during testing is output as a reference model, and a target model is selected from the multiple reference models based on the accuracy and recall rates of the multiple reference models, and further, the accuracy and recall rates corresponding to each defect of each reference model are obtained, and a confusion matrix of each reference model is obtained based on the accuracy and recall rates corresponding to each defect, and the target model is obtained based on the confusion matrix. When determining the target model, it is also possible to obtain the F1 score of each reference model and obtain the target model by referring to the F1 score and the confusion matrix, and this embodiment is not particularly limited.
[0070] Specifically, an optimal reference model is selected as the target model based on the accuracy and recall rates of multiple reference models in the confusion matrix. For example, the reference model with the highest accuracy and recall rate is selected as the target model, but this embodiment is not limited to this.
[0071] In one embodiment of the present disclosure, the method may further include modifying the confidence level based on the confusion matrix, specifically, by analyzing the accuracy rate and recall rate of each defect in the confusion matrix in detail, and adjusting the confidence level after the model is online in association with specific service requirements, thereby adjusting the accuracy and recall rate of the live model for a specific defect. For example, the recall rate of PI800 in the current confusion matrix is 0.90, and the recall rate is generated with a default confidence level of 0.8. Since PI800 is an unimportant defect, moderate over-identification is acceptable. In order to improve the recall rate of this defect in the production line, the confidence level of PI800 is set to 0.6-0.7 when the model is online, so that the recall rate of PI800 during production is improved to 0.91-0.92, and the accuracy of PI800 during production is accordingly reduced to 0.01-0.02. The increase in the recall rate can reduce the workload of the plotter. Therefore, users can customize the confidence level of each defect by conducting detailed analysis of the confusion matrix before putting the model online according to the confusion matrix and production requirements.
[0072] In step S350, actual data of defective products corresponding to the sample data set is input into the target model to obtain defect information of the defective products.
[0073] In this embodiment, after obtaining the target model, actual data of the product corresponding to the sample data is transmitted to the target model, and the target model is used to obtain defect information of the product, where the actual data of the product includes data of the detected product corresponding to the product defect data in the sample data set.
[0074] The present disclosure also provides a model training method, as shown in FIG. 5, comprising the following steps: In step S510, a sample data set including defective product data is obtained, and characteristic information of the sample data set is identified, the characteristic information including a number of samples of the sample data set; In step S520, an initial model is obtained, and the initial model is a neural network model; In step S530, training parameters are configured based on the characteristic information; In step S540, the initial model is trained based on the training parameters using the sample data set to obtain a target model, and the target model is used to perform defect detection on actual data of a product corresponding to the sample data set; The training parameters include at least one of a learning rate descent policy, a total number of training rounds, and a testing policy, the learning rate descent policy includes a number of times the learning rate is descent and a number of rounds during the descent, and the testing policy includes a number of times testing is performed and a number of rounds during testing.
[0075] In step S510, a sample dataset is obtained, and characteristic information of the sample dataset is identified, where the characteristic information includes the number of samples in the sample dataset.
[0076] In one embodiment of the present disclosure, in this embodiment, the feature information of the above sample data set includes the type, size, etc. of the picture of the defective product in the sample data, or includes the number of samples of the above sample data, where the number of samples may be 10,000, 20,000, etc., and may be defined according to the user's requirements, but is not limited to this in this embodiment.
[0077] In step S520, an initial model is obtained, and the initial model is a neural network model; In this embodiment, the initial model may be, but is not limited to, a convolutional neural network (CNN) model, a target detection convolutional neural network (faster-RCNN) model, a recurrent neural network (RNN) model, a generative adversarial network (GAN) model, or other neural network models known to those skilled in the art.
[0078] In this embodiment, the above initial model can be determined according to the type of defect picture in the above defective product.Specifically, in this embodiment, according to actual service requirements, three types of images may be involved: final site image (SD_final site), intermediate site image (mask site), and tdi grayscale image, and different initial models may be selected according to different images.For example, when the image of the sample data set is an intermediate site image, the initial model may be a convolutional neural network (CNN) model; when the image of the sample data set is a final site image, the initial model may be a convolutional neural network (CNN) model, a target detection convolutional neural network (faster-RCNN) model, but is not particularly limited in this embodiment.
[0079] In step S530, training parameters are configured based on the feature information.
[0080] In this embodiment, the training parameters include a learning rate descent policy, a total number of training rounds, and a test policy, where configuring the training parameters based on the above feature information may be configuring a learning rate descent policy, a total number of training rounds, and a test policy based on the number of samples in the feature information, where the learning rate descent policy includes the number of learning rate descents and the number of rounds during descent, and the test policy includes the number of tests and the number of rounds during testing.
[0081] Specifically, the total number of rounds of the training is positively correlated with the number of samples. For example, when the number of samples is less than or equal to 10000, the total number of rounds of the training is configured as 300000. When the number of samples is greater than 10000, the total number of rounds of the training is configured using the following formula for determining the total number of rounds, where the formula for determining the total number of rounds is Y=300000+INT(X / 10000)×b, where Y indicates the total number of rounds of the training, X indicates the number of samples, X is greater than or equal to 10000, INT is a rounding function, b is a growth factor and a fixed value, and b is greater than or equal to 30000 and less than or equal to 70000. In this embodiment, the value of b may be 50000 or 60000, but is not limited to this in this embodiment. In this embodiment, the mapping relationship between the number of samples and the total number of rounds of the training may be the result of optimization obtained by multiple tests, or may be defined according to the user's request, but is not limited to this in this embodiment.
[0082] In this embodiment, the number of rounds during the learning rate drop is positively correlated with the total number of training rounds, where the number of rounds during testing is equal to or greater than the number of rounds during the first learning rate drop and is equal to or less than the total number of training rounds, the number of times the learning rate is dropped is multiple, and at least two tests are performed within a predetermined range of the number of rounds of the second learning rate drop, and two tests may be performed, or three or more tests may be performed, but this embodiment is not limited thereto. Multiple learning rate drops are performed during training, and the number of drops with the best results after multiple drops is selected to improve the accuracy of the obtained target model, and further improve the accuracy of defect detection, and multiple tests are performed on the model during the training process, and the model with the best test result is selected as the target model, and further improve the accuracy of defect detection.
[0083] In this embodiment, the above learning rate drop method may be a division constant drop, an exponential drop, a natural exponential drop, a cosine drop, etc., and is not limited thereto in this embodiment; here, the above learning rate drop width is related to the above learning rate drop and is related to each parameter in the configured above learning rate drop method, and may be directly set to, for example, 0.1, 0.05, etc., and is not particularly limited thereto in this embodiment.
[0084] For details, please refer to the configuration method of the defect detection method described above, and the description will be omitted here.
[0085] In step S540, the sample data set is used to train the initial model based on the training parameters to obtain a target model, and the target model is used to perform defect detection on actual data of a product corresponding to the sample data set.
[0086] In one embodiment of the present disclosure, after configuration and modification to the above training parameters, the above obtained initial model can be trained using the above sample dataset to obtain a target model.
[0087] The target model is a neural network model that is primarily based on deep learning. For example, the target model may be based on a feedforward neural network. A feedforward network can be implemented as a ringless diagram with nodes arranged in layers. Typically, a feedforward network topology includes an input layer and an output layer separated by at least one hidden layer. The hidden layer converts the input received by the input layer into a representation that helps generate the output at the output layer. The network nodes are fully connected to the nodes of adjacent layers through edges, but there are no edges between the nodes within each layer. Data received at the nodes of the input layer of the feedforward network are propagated (i.e., "feed forward") to the nodes of the output layer through an activation function that calculates the state of the nodes of each successive layer of the network based on coefficients ("weights"), respectively associated with each edge connecting these layers. The output of the target model can take various forms that are not limited by this disclosure. The target model may also include other neural network models, such as, but not limited to, a convolutional neural network (CNN) model, a recurrent neural network (RNN) model, a generative adversarial network (GAN) model, or other neural network models known to those skilled in the art.
[0088] The training with sample data may include the steps of selecting a network topology, using a training data set representative of the problem to be modeled by the network, and adjusting the weights until the network model operates with minimal error for all instances of the training data set. For example, in the training process of supervised learning of a neural network, the outputs generated by the network in response to inputs representing instances of the training data set are compared with outputs marking the instances as "correct", an error signal representing the difference between the outputs and the marked outputs is calculated, and the weights associated with the connections are adjusted to minimize the error as the error signal is propagated backwards through the layers of the network. When the error for each output generated from an instance of the training data set is minimized, the initial model is considered "trained" and is defined as the target model.
[0089] In this embodiment, during the training of the initial model, the defect detection method further includes obtaining a loss curve during the training process, and then the training parameters may be adjusted according to the loss curve. Specifically, referring to FIG. 4, the abscissa of the loss curve is the number of training rounds, and the ordinate is the loss value. During the model training process, the loss curve is updated in real time according to the state during training, and the user can observe the loss curve and adjust the training parameters according to the state. Specifically, if the loss curve is always disturbed and does not show a downward trend, it is determined that the training parameters are not properly configured, and the training should be stopped, and the learning rate parameter and the learning rate descent policy should be readjusted and retrained. If the loss curve gradually drops, it should be continuously observed and the training should be stopped or the initial learning rate should be increased in the next training. If the loss curve still shows a downward trend after the training is completed (normally it should be smoothed out eventually), the maximum number of training rounds should be increased and retraining should be performed after the training is completed.
[0090] In one embodiment of the present disclosure, according to the test policy, a model that has reached the number of training rounds and the number of rounds during testing is output as a reference model, and a target model is selected from the multiple reference models based on the accuracy and recall rates of the multiple reference models, and further, the accuracy and recall rates corresponding to each defect of each reference model are obtained, and a confusion matrix of each reference model is obtained based on the accuracy and recall rates corresponding to each defect, and the target model is obtained based on the confusion matrix. When determining the target model, it is also possible to obtain the F1 score of each reference model and obtain the target model by referring to the F1 score and the confusion matrix, and this embodiment is not particularly limited.
[0091] Specifically, an optimal reference model is selected as the target model based on the accuracy rate and recall rate of multiple reference models in the confusion matrix, for example, the reference model with the highest accuracy rate and recall rate is selected as the target model, but this embodiment is not limited thereto. For details of the training of the initial model, please refer to the training of the initial model in the defect detection method, and the description will be omitted here.
[0092] The present disclosure also provides a model training method, as shown in FIG. 6 , the model training method includes the following steps: In step S610, in response to a user's configuration operation on parameters of the sample data set, a sample data set including defective product data is obtained, and characteristic information of the sample data set is identified, the characteristic information including a number of samples of the sample data set; In step S620, an initial model is obtained, the initial model being a neural network model; In step S630, configure training parameters based on the feature information, and generate a display interface for the training parameters; In step S640, the sample data set is used to train the initial model based on the training parameters to obtain a target model, and the target model is used to perform defect detection on actual data of a product corresponding to the sample data set.
[0093] The training parameters include at least one of a learning rate descent policy, a total number of training rounds, and a testing policy, the learning rate descent policy includes a number of times the learning rate is descent and a number of rounds during the descent, and the testing policy includes a number of times testing is performed and a number of rounds during testing.
[0094] Each of the above steps will now be described in detail.
[0095] In step S610, a sample dataset including defective product data is obtained in response to a user's configuration operation on parameters of the sample dataset, and characteristic information of the sample dataset is identified, the characteristic information including the number of samples of the sample dataset.
[0096] In one embodiment of the present disclosure, sample data is acquired in response to a user's configuration operation on parameters of a sample data set, for example, sample data corresponding to a number of defects is displayed in a graphic user interface, and an acquired icon corresponding to the sample data is displayed, and when a user triggers the acquisition icon, sample data corresponding to the icon is acquired.
[0097] In this embodiment, as shown in FIG. 20 , in response to a user's task establishment operation, establish a training task and generate a parameter configuration interface for the sample data set, where the user can configure parameters for the sample data set in the parameter configuration interface; Here, the parameters of the sample data set may include department, section, site, picture type, product, training type, etc., and in response to the user's parameter configuration operation, a sample data set corresponding to the parameters of the sample data can be automatically obtained, and feature information in the sample data can be identified. In another embodiment, after obtaining the sample data, a training task corresponding to the sample data set can be created based on the parameters of the sample data set, but this embodiment is not particularly limited.
[0098] For obtaining feature information for the above sample data set, please refer to the above content, and the explanation will be omitted here.
[0099] In step S620, an initial model is obtained, where the initial model is a neural network model.
[0100] In this embodiment, obtaining the initial model refers to the description of the defect detection method, and the description is omitted here.
[0101] In step S630, configure training parameters based on the feature information, and generate a display interface for the training parameters; In an embodiment of the present disclosure, the configuration of training parameters based on the feature information refers to the contents of the defect detection method described above, and a description thereof will be omitted here.
[0102] In one embodiment of the present disclosure, as shown in Fig. 20, the above sample configuration interface includes a training parameter view icon, and the server generates a display interface of the above training parameters in response to a user's trigger operation on the training parameter view icon. In another embodiment, as shown in Fig. 7, after the above training parameters are configured, a display interface of the training parameters is directly generated, where the display interface of the training parameters includes configuration information and a parameter change icon for each of the above training parameters.
[0103] Here, the above training parameters include the total number of training rounds, the learning rate descent policy, the testing policy, the confidence level, and the size of the image input to the initial model, and may include other parameters, but are not limited to these in this embodiment.
[0104] In this embodiment, the total number of training rounds, the learning rate descent policy, the test policy, the confidence level, and the size of the image input to the initial model refer to the description of the defect detection method above, and the description is omitted here.
[0105] In this embodiment, referring to Fig. 7, the server changes the training parameters in response to the user's trigger operation on the parameter change icon, and after triggering the change icon, each training parameter is in a changeable state and a confirmation icon is generated, and after the user triggers the confirmation icon, the change to the training parameters is completed. The training parameters are changed through an interaction interface, and there is no need to directly change the code, which facilitates the operation by system operation and maintenance personnel who do not understand the program, and improves convenience.
[0106] In one embodiment of the present disclosure, the feature information further includes a defect level of the above-mentioned defective product and the number of samples corresponding to various defects, and the training parameters further include a reliability, where the reliability in the training process can be configured based on the number of samples corresponding to various defects and the defect level, and where the display interface of the above-mentioned parameters further includes a reliability configuration icon.
[0107] In this embodiment, as shown in Fig. 8, in response to a trigger operation by a user on a reliability configuration icon, a reliability configuration interface is generated, where the reliability configuration interface includes a sample number corresponding to each defect and a selection icon corresponding to each defect, and the reliability configuration interface is used for a user to respond to a decision operation on the selection icon to configure a reliability of the defect corresponding to the selection operation, specifically, in response to a decision operation by a user on a selection icon whose sample number in the reliability configuration interface is greater than a predetermined number of defects, a reliability is configured based on a defect level, where the defect level includes a first defect level and a second defect level, and when the defect level of the defect is the first defect level, the reliability is configured as a first reliability level, and when the defect level of the defect is the second defect level, the reliability is configured as a second reliability level, where the second reliability level is greater than the first reliability level.
[0108] In this embodiment, the above-mentioned predetermined number is 50, 100, etc., and may be defined according to the user's request, but is not limited thereto in this embodiment. Here, the first reliability is 0.6 or more and 0.7 or less, and the second reliability is 0.8 or more and 0.9 or less, and the above-mentioned first reliability and second reliability values may be defined according to the user's request, but is not limited thereto in this embodiment.
[0109] For example, for defects with a high occurrence rate and low importance, i.e., defects with a low defect level, a low reliability can be set, for example, for PI820 with no defect and PI800 with a minor defect, the reliability is set to 0.6, i.e., in this figure, it can be determined that there is a defect if the probability score for PI800 or PI820 exceeds 0.6. For defects with a low occurrence rate and high importance, i.e., defects with a high defect level, a higher reliability can be set, for example, for GT011 and SD011, which are serious defects, the reliability is set to 0.85, i.e., in this figure, it is determined that there is a defect only if the probability score for GT011 or SD011 exceeds 0.6. For the other figures with low reliability, all are determined to be unrecognized, and manual processing is performed to prevent missed determinations.
[0110] In step S640, the sample data set is used to train the initial model based on the training parameters to obtain a target model, and the target model is used to perform defect detection on actual data of a product corresponding to the sample data set.
[0111] In this embodiment, the training process refers to the above description of the defect detection method, and the description is omitted here.
[0112] In one embodiment of the present disclosure, as shown in FIG. 9 , when training the above initial model, a training schedule is generated and displayed, where the above training schedule includes a task details icon and a task cancel icon, and when a user triggers the task details icon, a loss curve in the training process is generated and displayed, and then the user can adjust the above training parameters according to the loss curve.
[0113] Adjusting the training parameters according to the loss curve refers to the above defect detection method description, and is not described here. In this embodiment, when the user triggers the task cancel icon, the training for the initial model is stopped.
[0114] In one embodiment of the present disclosure, as shown in FIG. In one embodiment of the present disclosure, according to the test policy, a model that has reached the number of training rounds and reached the number of rounds during testing is output as a reference model, and a target model is selected from the multiple reference models based on the accuracy and recall rates of the multiple reference models, and further, the accuracy and recall rates corresponding to each defect of each reference model are obtained, and a confusion matrix of each reference model is obtained based on the accuracy and recall rates corresponding to each defect, and the target model is obtained based on the confusion matrix, and a display interface of the confusion matrix is generated, and the target model is obtained based on the confusion matrix. When determining the target model, it is also possible to obtain the F1 score of each reference model and obtain the target model by referring to the F1 score and the confusion matrix, and this embodiment is not particularly limited.
[0115] Specifically, as shown in Fig. 11, the optimum reference model is selected as the target model based on the accuracy rate and recall rate of the multiple reference models in the confusion matrix, specifically, the user responds to the selection operation for the multiple reference models, and the reference model corresponding to the selection operation is selected as the target model, for example, the reference model with the maximum accuracy rate and recall rate is selected as the target model, which is not limited in this embodiment. After the selection, the user clicks the determination icon to determine the target model.
[0116] In one embodiment of the present disclosure, the user can update the confidence level based on the confusion matrix in response to the user's change operation on the confidence level, Specifically, the accuracy rate and recall rate of each defect in the confusion matrix may be analyzed in detail, and the reliability of the model after it is online may be adjusted in association with specific service requirements, thereby adjusting the accuracy and recall rate of the live model for a specific defect. For example, the recall rate of PI800 in the current confusion matrix is 0.90, and the recall rate is generated with a default reliability of 0.8. Since PI800 is an unimportant defect, moderate over-identification is acceptable. In order to improve the recall rate of this defect in the production line, the reliability of PI800 is set to 0.6-0.7 when the model is online, so that the recall rate of PI800 during production is improved to 0.91-0.92, and the accuracy of PI800 during production is accordingly reduced to 0.01-0.02. The increased recall rate can reduce the workload of the plotter. Therefore, the user can customize the reliability of each defect by performing a detailed analysis of the confusion matrix before the model is online according to the confusion matrix and production requirements.
[0117] Further, the present disclosure also provides a defect detection system, and as shown in FIG. 12 , said system includes a data management module 1202, a model management module 1204 and a training management module 1203, where the data management module 1202 is configured to store and manage sample data, the training management module 1203 is configured to execute the above defect detection method and model training method, and the model management module 1204 is configured to store, display and manage the target model.
[0118] In one embodiment of the present disclosure, the above defect detection system further includes a user management module, where the user management module 1201 is configured to add, delete, modify, view, manage permissions and / or manage passwords for user information.
[0119] Referring to Fig. 13, the above system adopts a BS or CS architecture, and is composed of a back-end server 1306, a front-end server 1303, and a shared storage 1304, and can be operated by a browser using a factory PC 1307 on the operator's side. As a subsystem responsible for training-related tasks in the ADC system, the training system communicates with the data mark system, the TMS system, and the GPU server (algorithm system) of the ADC system. The training system is connected to the data mark system and the TMS system, and can provide the TMS system 1305 with updated target models and other related services.
[0120] The training system interacts with the data mark system 1302 for data and images through the database and shared storage 1304 (NAS network storage), the training system interacts with the GPU server (algorithm system) 1301 to communicate through TCP / IP protocol, thus controlling the GPU server 1301 to perform model training and automatic testing, the training system transfers model data with the TMS system 1305 through FTP protocol, and interacts with model information through the database. The training system uses HTTP protocol to interact with front-end and back-end services and web interfaces.
[0121] Specifically, the user management module is the user management and system information management module of the system, which is used for the functions of adding, deleting, modifying and checking user information, authority management, password management, adding, deleting, modifying and checking work divisions, departments and sites applied to the system. The user management module can include user information and system information, and the user can enter the training system by entering the user name and password, and then enter the completed training module by default, which is convenient for the user to directly view the training status of the existing model. All users currently managed can be checked, and authority management can also be set in the system, which can add, modify and delete users.
[0122] In an exemplary embodiment of the present disclosure, the data management module is configured to store and manage sample data, and since the deep learning AI algorithm is a data-driven approach, differences in the production process and the AOI capture device in the factory will lead to large differences in and between data types, making it difficult to use one general-purpose model to solve all problems. Therefore, a model can be individually trained for specific data to comprehensively cover real-time inference services in factory production.
[0123] This dataset management module can process these datasets in a unified manner, which makes it easier to train models. The dataset management is divided into training dataset management and preparatory dataset management, where the data in the preparatory dataset management is the original data marked by the data labeling system and can be imported into the training data management after statistical verification by the user, and the data in the training dataset management can be directly sent to the corresponding training task to perform model training.
[0124] Specifically, as shown in Figure 14, the data marked by the system is automatically synchronized (manual synchronization is also possible) to the Data Set Management, and the original data set synchronization and statistical display function can be performed in the Data Set Management. In the Data Set Management, details such as the product corresponding to each data set, the number of defect types, the number of pictures, the minimum number of defects in a picture, and the update time are mainly displayed.
[0125] In Dataset Management, users can click the Sync icon to directly sync the data marked by Datamark system, and the system can synchronize data on a daily basis without manual synchronization. Users can also click the Statistics icon to get detailed information for each dataset, such as the types of defects contained in each dataset, the number of pictures of each type, and the distribution table.
[0126] In managing the training dataset, referring to FIG. 15, the above preparation dataset data is statistically checked, and then functions such as row creation, modification, and picture management are performed to perform a series of operations on the dataset to generate a sample dataset that can be used to train a model.
[0127] There is a correspondence between sample datasets and models, and each sample dataset can be trained with different training parameters to train different models, and users can select models according to related metrics. Therefore, the first step of preparation model training is to create a new sample dataset that corresponds to the model to be trained in this interface.
[0128] There are specific rules for the name of the sample data set of the defect detection system. In other words, the training parameters for the system to train can be automatically set based on the feature information of the sample data set. Since the name of the sample data set is also a feature of the data set, the naming of the sample data set must follow specific rules. For example, in the case of the SD / Final site model, the name of the Sub-defect model can be "product_product name" and the name of the Main-defect model can be "main-defect_product name" (where the product name is two letters or numbers). For example, the name of the sub-defect data set of ak is product_ak, and the name of the main-defect data set of ak is main-defect_ak.
[0129] In this embodiment, for the Mask site model, the sub defect model of the Mask site is "mask_site name" (the site name is four letters or numbers), and the Mask site does not have a main defect model. For example, the model data set of 1500 sites is named mask_1500.
[0130] In this embodiment, the TDI model is currently applied to the SD / Final site and is not a main defect model. The naming rule is "tdi_tdiX" (X is a number). For example, the model name corresponding to the conventional product is tdi_tdi1.
[0131] As shown in Figure 16, after selecting the department, section, and site, the user needs to manually input the name of the sample data set into the system according to the naming rules mentioned above, and when the input is completed, a new data set is completed. After clicking the above change icon, enter a new data set name and click Confirm to complete the data set name change, and click the delete icon after the corresponding sample data set to complete the deletion of the sample data set, where the deletion deletes the data imported into this module, and as shown in Figure 17, clicking the details icon above displays detailed information in each sample data set, including the location of the sample data set, the internal organizational structure, and picture preview. Here, the center column is the directory structure of pictures in the sample data set, and the rightmost column is the picture list in the selected defect category folder, where pictures can be previewed and deleted.
[0132] The model management module is configured to store, display and manage the target models, and will be specifically described with reference to Fig. 18. Including renaming, deleting and displaying details of the target model, when a user clicks the details icon, all version information of a certain model will be displayed, including model name, model version number, whether the model is online, trainer, training completion time, model accuracy rate, recall rate, model path, etc. Click the delete icon to delete a certain trained target model.
[0133] The training management module realizes the submission of training tasks, adaptive configuration of training parameters, management of training tasks, and display of automatic training test results, where adaptive configuration of training parameters and automatic training test are functions not found in traditional training systems. These two functions greatly reduce the difficulty of algorithm model training, and write the parameter adjustment experience performed by the algorithm user during development into the system logic, realizing adaptive parameter adjustment and automatic testing, so that even operation and maintenance personnel and users who do not develop algorithms can use this module to train models that can achieve the accuracy of the production line. The training management module is divided into three sub-modules: trained, training, and cancel.
[0134] In this embodiment, refer to Fig. 19. A training task can be submitted by clicking the new icon on any of the screens of "Trained", "Training", and "Cancel". Note that a training task can be submitted only if a corresponding sample dataset has been created or if a sample dataset related to the training task is directly obtained from the system according to the training task.
[0135] To facilitate model training, for the complex production scenario in which the defect detection system is located, and the data characteristics of large intra-type and inter-type differences, the training system has the function of automatically setting training parameters based on the characteristics of the data set, as shown in Figure 20. After selecting "Site", "Picture Type", "Product" (see the table below for the product naming rules) and "Training Type", click "Training Parameters" to confirm the changes, wait for a few seconds, and the relevant configuration parameters of the model will automatically pop up, for details, please refer to the contents shown in Figure 7.
[0136] [Table 1]
[0137] For the naming rules of the product name, refer to Table 1, select the relevant dataset, initial model (optional), input model name (optional), set the training type to mainly main defects and sub defects (the corresponding relationship between training type and picture type is shown in Table 2), change the training parameters, and then click Confirm.
[0138] [Table 2]
[0139] 7 and 8, the defect configuration icon on the interface needs to be clicked, and in the pop-up interface, check according to the picture of the data set and change the confidence level as necessary. The specific check is that a predetermined number or more of samples are selected, where the predetermined number is described in detail in the above defect detection method, so the description is omitted here.
[0140] In this implementation, when a page like that shown in Figure 8 pops up, the system gives default values for individual defect confidence levels based on the algorithm user's experience, and these default values are determined by multiple tests based on the importance and occurrence rate of defects provided by the service. For certain defects with high occurrence rates and low importance, the confidence level is relaxed to a certain extent.
[0141] For example, for PI820 with no defect and PI800 with minor defect, the reliability is set to 0.6, i.e., in this figure, if the probability score for PI800 or PI820 exceeds 0.6, it can be determined that there is a defect. For defects with low occurrence rate and high importance, the reliability is treated strictly, for example, for GT011 and SD011, which are serious defects, the reliability is set to 0.85, i.e., in this figure, only when the probability score of GT011 or SD011 exceeds 0.6, it is determined that there is a defect. For other figures with low reliability, all are determined to be unknow (not recognized by AI), and manual processing is performed to prevent missed judgments.
[0142] As mentioned above, for every defect that pops up, the optimal confidence level is selected by the algorithm user through several tests based on the requirements of the service, and the default setting is made. During the training process, the above confidence level configuration conditions can be developed autonomously.
[0143] Referring to FIG. 9, in the training management module, the training task can be checked, including information such as the dataset name of the training task, the trainer, the reason for training, the current training round number, and the total number of training rounds. There are two operation icons, Cancel and Details, for operations. Clicking the Cancel icon will cancel the current training task, and clicking the Details icon will generate and display the training loss curve of the current training task, as shown in FIG. 4. The training effect of the model can be judged to some extent by the trend of the loss curve.
[0144] In this embodiment, the abscissa of the loss curve is the number of training rounds, and the ordinate is the loss value. During the model training process, the loss curve is updated in real time according to the state during training, and the user can observe the loss curve and adjust the training parameters according to the state. Specifically, if the loss curve is always disturbed and does not show a downward trend, it is determined that the training parameters are not properly configured, and training should be stopped, and the learning rate parameter and the learning rate descent policy should be readjusted and retrained. If the loss curve gradually drops, it should be continuously observed and training should be stopped or the initial learning rate should be increased in the next training. If the loss curve still shows a downward trend after the training is completed (normally it should eventually smooth out), the maximum number of training rounds should be increased and retraining should be performed after the training is completed.
[0145] Principle of confidence setting: In the confidence input box of the training parameter interface, the default confidence for all defects can be set. If the confidence for each defect in the training defect and confidence setting interface is not selected and entered, the default value on the previous page will be used by default; if a value is entered, the confidence in the interface will be used.
[0146] Referring to FIG. 21, the training management module displays the trained models, displays the models that have completed training, including the name of the model generated for each dataset (the models are displayed here as one dataset corresponds to models with different training round numbers, and instead of one model, the model selected by the test when submitting for training is displayed here), the time when training was completed, and the reason for training.
[0147] Referring to Figure 19, the operation includes two icons, Check Results and Select Optimal, which are Check Results and Select Optimal, respectively. Clicking Check Results will pop up a confusion matrix of multiple (related to the number of test rounds set at the beginning of the training parameter configuration, the default is to test six models) models. This function complements the traditional training system, converting the indicators related to the algorithm into a more intuitive table to display the data, allowing operation and maintenance personnel and users to view the data in a familiar table format and select a model based on the indicators of interest. Clicking the Select Optimal button will display the test results of each test model, where the main indicators are accuracy, recall, and F1 score, and based on the accuracy, recall, F1 score, and the confusion matrix mentioned above, select whether to put the model online.
[0148] In addition to the role of guiding the model online, the user can also perform detailed analysis of the accuracy rate and recall rate of each defect in the confusion matrix, and adjust the reliability after the model is online, taking into account specific service requirements, to adjust the accuracy rate and recall rate of the online model for specific defects. For example, the recall rate of PI800 in the current confusion matrix is 0.90, and the recall rate is generated with a default reliability of 0.8. Since PI800 is an insignificant defect, moderate overjudgment is acceptable. In order to improve the recall rate of this defect in the production line, the reliability of PI800 is set to 0.6-0.7 when the model is online, so that the recall rate of PI800 during production is improved to 0.91-0.92, and the accuracy of PI800 during production is accordingly reduced to 0.01-0.02. The increase in the recall rate can reduce the workload of the plotter. Therefore, the user can customize the reliability of each defect by performing detailed analysis of the confusion matrix before the model is online according to the confusion matrix and production requirements. After selecting check model, check this model in the model management interface of TMS system and run it on the production line (it can also be checked in the model management function module).
[0149] It should be noted that the above-mentioned attached drawings are merely schematic illustrations of steps included in the method according to the exemplary embodiment of the present disclosure, and are not intended to be limiting. It is easily understood that the processes illustrated in the above-mentioned attached drawings do not indicate or limit the time order of the processes. It is also easily understood that these processes may be performed, for example, synchronously or asynchronously in multiple modules.
[0150] In this embodiment, referring to FIG. 22, the sample data set generated by the data mark system 2201 is used to call the defect detection system 2202 of the present disclosure, which is trained into a corresponding model, and uploaded to the TMS system 2203 for model inference service to be online. The specific process is described below, and the user organizes and marks the training data in the data mark system. With the embedded AI supervision algorithm and the traditional unsupervised algorithm, the automatic data mark function and the semi-automatic data mark function are implemented, which greatly reduces the workload of data marking. The data marked in this module is automatically synchronized to the relevant directory in the training management module. Import the training sample data set in the training management module. Create a new sample data set name according to the training task, import it into the data mark system to mark the completed data, and write the relevant information to the database. After the data is imported, operations such as detailed confirmation, data modification, and deletion can be performed. Submit the training task. After the training set is imported, the corresponding training task is submitted in the training task management module. In the process of submitting a training task, the user needs to select information such as site, picture type, training type, etc. according to the characteristics of the sample dataset. After selecting the relevant information, the training management module will adaptively adjust the training parameters, and the configured training parameters will be displayed on the interface for the user's reference, and the user can choose to use or change the default parameters. Once completed, the training task can be submitted. Perform automatic training and testing of the model. After submitting the training task, the background of the training management module will send the training task and the configured training parameters to the AI algorithm system. Then, the AI algorithm system will automatically train the model according to the received task, and automatically test the model after training is completed.Various indicators that change during the training process are displayed as icons on the system, so that users can grasp the training status. Processing the results of the training test: After the training is completed, the AI algorithm system stores the target model in the shared storage and sends the related training results to the training management module. After receiving the message from the AI algorithm system, the training management module displays the training results in the trained section, and the user can view the automatically tested indicators of the model, the confusion matrix, and some necessary training indicator analysis. Making the model online: Based on the above indicators, the user can select the optimal model and make it online, and synchronize the model to the model management module of the TMS system 2203. The pre-release model can be easily viewed and managed in the training management module. Making the model online: The pre-release model has been synchronized to the model management database of the TMS system 2203, and the user can test the model offline and make it online in the TMS system.
[0151] Further, as shown in FIG. 23, this embodiment provides a defect detection apparatus 2300 , which includes a first acquisition module 2310 , a first configuration module 2320 , a first training module 2330 and a detection module 2340 . Here, the first acquisition module 2310 is configured to acquire a sample dataset including defective product data, and identify feature information of the sample dataset, where the feature information includes the number of samples of the sample dataset, and is configured to acquire an initial model, where the initial model is a neural network model, the first configuration module 2320 is configured to train parameters based on the feature information, the first training module 2330 is configured to use the sample dataset and train the initial model based on the training parameters to obtain a target model, and the detection module 2340 is configured to input actual data of the product corresponding to the sample dataset into the target model to obtain defect information of the product, where the training parameters include at least one of a learning rate descent policy, a total number of training rounds, and a test policy, where the learning rate descent policy includes the number of learning rate descents and the number of rounds during descent, and the test policy includes the number of tests and the number of rounds during testing.
[0152] Further, as shown in Fig. 24, this embodiment provides a model training apparatus 2400, which includes a second acquisition module 2410, a second configuration module 2420, and a second training module 2430. Here, the second acquisition module 2410 is configured to acquire a sample dataset including defective product data, identify feature information of the sample dataset, the feature information includes the number of samples of the sample dataset, and is configured to acquire an initial model, the initial model is a neural network model, the second configuration module 2420 is configured to configure training parameters based on the feature information, the second training module 2430 is configured to use the sample dataset to train the initial model based on the training parameters to obtain a target model, the target model is used to perform defect detection on real data of a product corresponding to the sample dataset, the training parameters include at least one of a learning rate descent policy, a total number of training rounds, and a test policy, the learning rate descent policy includes the number of descents of the learning rate and the number of rounds during descent, and the test policy includes the number of tests and the number of rounds during testing.
[0153] Further, as shown in FIG. 25, this embodiment provides a model training apparatus 2500, which includes a third acquisition module 2510, a third construction module 2520, and a third training module 2530. Here, the third acquisition module 2510 is configured to acquire a sample dataset including defective product data in response to a user's configuration operation on parameters of the sample dataset, and identify feature information of the sample dataset, where the feature information includes the number of samples in the sample dataset, and to acquire an initial model, where the initial model is a neural network model; the third configuration module 2520 is configured to configure training parameters based on the feature information and generate a display interface for the training parameters; the third training module 2530 is configured to use the sample dataset to train the initial model based on the training parameters to obtain a target model, where the target model is used to perform defect detection on actual data of a product corresponding to the sample dataset; the training parameters displayed in the display interface for the training parameters include at least one of a learning rate descent policy, a total number of training rounds, and a test policy, where the learning rate descent policy includes the number of descents of the learning rate and the number of rounds during descent, and the test policy includes the number of tests and the number of rounds during testing.
[0154] The specific contents of the modules in the above-mentioned device are described in detail in the method section implementation contents, and since the contents not disclosed can be confirmed in the method section implementation contents, the explanation thereof will be omitted here.
[0155] Those skilled in the art will appreciate that aspects of the present disclosure may be embodied as a system, method, or program product. As such, aspects of the present disclosure may be specifically implemented in an entirely hardware implementation, an entirely software implementation (including firmware, microcode, etc.), or a combination of hardware and software aspects, collectively referred to herein as a "circuit," "module," or "system."
[0156] Exemplary embodiments of the present disclosure also provide a computer-readable storage medium on which a program product capable of implementing the above-mentioned methods of the present disclosure is stored. In some possible implementations, various aspects of the present disclosure may also be implemented in the form of a program product including program code, and when the program product is executed on a terminal device, the program code is used to execute the terminal device. The above-mentioned steps of the procedure according to various exemplary embodiments of the present disclosure are described in the "Example Method" section.
[0157] It should be noted that the computer-readable medium described in this disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, apparatus, or any combination thereof. More specific examples of the computer-readable storage medium may be an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic memory device, or any suitable combination of the above.
[0158] In this disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, device, or apparatus. And, for purposes of this disclosure, a computer-readable signal medium may include a data signal propagating in baseband or as part of a carrier wave that carries computer-readable program code. Such propagated data signals may take various forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination of the foregoing. Also, a computer-readable signal medium may be any computer-readable medium other than a computer-readable storage medium that transmits, propagates, or transmits a program for use by an instruction execution system, device, or apparatus. The program code contained in the computer-readable medium may be transmitted using any suitable medium, such as wireless, wired, fiber optic cable, RF, etc., or any suitable combination of the foregoing.
[0159] Additionally, program code for carrying out the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, and the like, and conventional procedural programming languages such as "C." The program code may be executed entirely on the user computing device, partially on the user device, as a standalone package, partially on the user computing device, partially on a remote computing device, or entirely on a remote computing device or server. If a remote computing device is included, the remote computing device may be connected to the user computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., connected via the Internet using an Internet Service Provider).
[0160] Other embodiments of the present disclosure will be readily apparent to those skilled in the art from consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure, including those means generally known or customary in the art but not disclosed herein, in accordance with the general principles of the present disclosure. The description and embodiments are considered to be exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.
[0161] It should be understood that the present disclosure is not limited to the exact configuration already described above and illustrated in the accompanying drawings, but various modifications and changes can be made without departing from the scope thereof, which is limited only by the appended claims. [Explanation of symbols]
[0162] 100 System Configuration 101, 102, 103 Terminal Devices 104 Network 105 Server
Claims
1. Obtaining a sample data set including defective product data and identifying characteristic information of the sample data set; Obtaining an initial model, configuring training parameters based on the characteristic information; training the initial model using the sample data set based on the training parameters to obtain a target model; inputting actual data of a product corresponding to the sample data set into the target model to obtain defect information of the product; the feature information includes a number of samples in the sample data set; the initial model is a neural network model; The training parameters include at least one of a learning rate descent policy, a total number of training rounds, and a test policy, the learning rate descent policy including the number of times the learning rate is descent and the number of rounds during descent, and the test policy including the number of tests and the number of rounds during testing. A defect detection method comprising:
2. The total number of rounds of training is positively correlated to the number of samples.
2. The defect detection method according to claim 1.
3. Configuring a total number of rounds of training according to a predetermined rule based on the number of samples If the number of samples is less than or equal to 10,000, configuring the total number of rounds of training to be 300,000; If the number of samples is greater than 10,000, configuring the total number of rounds of the training with the following formula for determining the total number of rounds, wherein the formula for determining the total number of rounds is: Y=300000+INT(X / 10000)×b, Y denotes the total number of training rounds, X denotes the number of samples, and X is equal to or greater than 10,000. INT is a rounding function, b is a growth factor and is a fixed value, and b is equal to or greater than 30,000 and equal to or less than 70,000.
3. The defect detection method according to claim 1 or 2.
4. The number of rounds during the learning rate reduction is positively correlated with the total number of rounds in the training, and the number of rounds during the test is equal to or greater than the number of rounds during the first learning rate reduction and equal to or less than the total number of rounds in the training.
2. The defect detection method according to claim 1.
5. The learning rate is lowered multiple times, and at least two tests are performed within a predetermined range of the number of rounds of the second learning rate descent.
2. The defect detection method according to claim 1.
6. The number of times the learning rate is decreased is three, and at least three tests are performed within a predetermined range of the number of rounds of the second learning rate decrease.
6. The defect detection method according to claim 5.
7. The learning rate reduction policy includes a learning rate reduction method and a learning rate reduction range.
2. The defect detection method according to claim 1.
8. The defective product data includes a picture of a defective product, and the characteristic information includes a size and a type of the picture of the defective product in the sample data set, and configuring training parameters based on the characteristic information further includes: and adjusting the size of the input image of the initial model according to the size and type of the picture of the defective product.
2. The defect detection method according to claim 1.
9. Adjusting the size of the image input to the initial model according to the size and type of the defective product picture; If the type of the defective product picture is an AOI color image or a DM image, the size of the input image is a first preset multiple of the size of the defective product picture; if the type of the defective product picture is a TDI image, the size of the input image is a second preset multiple of the size of the defective product picture; The first preset multiple is equal to or less than 1, and the second preset multiple is equal to or greater than 1.
9. The defect detection method according to claim 8.
10. The first preset multiple is equal to or greater than 0.25 and equal to or less than 0.6, and the second preset multiple is equal to or greater than 3 and equal to or less than 6.
10. The defect detection method according to claim 9.
11. The input images include multiple sized images of the same picture of the defective product.
11. The defect detection method according to claim 9 or 10.
12. the characteristic information further includes a defect level of the defective product; Configuring training parameters based on the feature information includes: and constructing a confidence level within the training process based on said defect levels corresponding to each type of said defect.
9. The defect detection method according to claim 8.
13. the defect levels include a first defect level and a second defect level; Constructing a confidence level within a training process based on the defect level corresponding to each type of defect, if the defect level is a first defect level, configuring the confidence level as a first confidence level; and if the defect level is a second defect level, configuring the confidence level as a second confidence level; The second reliability is greater than the first reliability. The defect detection method according to claim 12 .
14. The first reliability is 0.6 or more and 0.7 or less, and the second reliability is 0.8 or more and 0.9 or less.
14. The defect detection method according to claim 13.
15. The method further comprises: and setting an icon for the target model using the size, type and process of the defective product picture.
9. The defect detection method according to claim 8.
16. Obtaining the initial model based on a type of picture of the defective product.
9. The defect detection method according to claim 8.
17. After configuring training parameters based on the feature information, the method further comprises: generating a training parameter display interface including a parameter change icon; and updating the training parameters in response to a user's trigger operation on the parameter change icon.
2. The defect detection method according to claim 1.
18. The method further comprises: Obtaining a loss curve during the training process; updating the training parameters according to the loss curve.
2. The defect detection method according to claim 1.
19. Training the initial model based on the training parameters to obtain a target model includes: Obtaining a plurality of reference models based on the test policy, and obtaining precision rates and recall rates of the plurality of reference models; determining a target model from the reference models based on the precision and recall rates of each reference model.
2. The defect detection method according to claim 1.
20. Training the initial model based on the training parameters to obtain a target model includes: obtaining a plurality of reference models based on the test policy and determining a confusion matrix for each reference model; determining the target model from the reference model according to the confusion matrix.
2. The defect detection method according to claim 1.
21. The method further comprises: updating the confidences according to the confusion matrix.
21. The defect detection method according to claim 20.
22. Obtaining a sample data set including defective product data and identifying characteristic information of the sample data set; Obtaining an initial model, configuring training parameters based on the characteristic information; and training the initial model based on the training parameters using the sample data set to obtain a target model; the feature information includes a number of samples in the sample data set; the initial model is a neural network model; the target model is used to perform defect detection on actual data of a product corresponding to the sample data set; The training parameters include at least one of a learning rate descent policy, a total number of training rounds, and a test policy, the learning rate descent policy including the number of times the learning rate is descent and the number of rounds during descent, and the test policy including the number of tests and the number of rounds during testing. A model training method comprising:
23. The total number of rounds of training is positively correlated to the number of samples.
23. The method of claim 22, wherein the model training method
24. Configuring a total number of rounds of training according to a predetermined rule based on the number of samples If the number of samples is less than or equal to 10,000, configuring the total number of rounds of training to be 300,000; If the number of samples is greater than 10,000, configuring the total number of rounds of the training with the following formula for determining the total number of rounds, wherein the formula for determining the total number of rounds is: Y=300000+INT(X / 10000)×b, Y denotes the total number of training rounds, X denotes the number of samples, and X is equal to or greater than 10,000. INT is a rounding function, b is a growth factor and is a fixed value, and b is equal to or greater than 30,000 and equal to or less than 70,000.
24. A method for training a model according to claim 22 or 23.
25. The number of rounds during the learning rate reduction is positively correlated with the total number of rounds in the training, and the number of rounds during the test is equal to or greater than the number of rounds during the first learning rate reduction and equal to or less than the total number of rounds in the training.
23. The method of claim 22, wherein the model training method
26. The learning rate is lowered multiple times, and at least two tests are performed within a predetermined range of the number of rounds of the second learning rate descent.
23. The method of claim 22, wherein the model training method
27. The number of times the learning rate is decreased is three, and at least three tests are performed within a predetermined range of the number of rounds of the second learning rate decrease.
24. The method of claim 23,
28. The learning rate reduction policy includes a learning rate reduction method and a learning rate reduction range.
23. The method of claim 22, wherein the model training method
29. acquiring a sample data set including defective product data in response to a configuration operation by a user on parameters of the sample data set; and identifying characteristic information of the sample data set; Obtaining an initial model, configuring training parameters based on the characteristic information and generating a display interface for the training parameters; and training the initial model based on the training parameters using the sample data set to obtain a target model; the feature information includes a number of samples in the sample data set; the initial model is a neural network model; the target model is used to perform defect detection on actual data of a product corresponding to the sample data set; The training parameters displayed in the training parameter display interface include at least one of a learning rate descent policy, a total number of training rounds, and a test policy, the learning rate descent policy includes the number of times the learning rate is descent and the number of rounds during descent, and the test policy includes the number of tests and the number of rounds during testing. A model training method comprising:
30. The method further comprises: establishing a training task corresponding to the sample data set based on parameters of the sample data set.
30. The method of claim 29, wherein the model training method
31. Prior to acquiring the sample data set including defective product data in response to a user configuration operation on parameters of the sample data set, the method further comprises: and generating a parameter configuration interface for the sample data set in response to a task establishment operation by a user.
30. The method of claim 29, wherein the model training method
32. The parameter configuration interface further includes a training parameter viewing icon, and configuring training parameters based on the characteristic information and generating a display interface for the training parameters includes: and generating a training parameter display interface in response to a user's trigger operation on the training parameter viewing icon.
32. The method of claim 31 .
33. The training parameter display interface includes a parameter change icon, and the method further includes: updating the training parameters in response to a trigger operation on the parameter change icon by a user.
30. The method of claim 29, wherein the model training method
34. The training parameters further include a confidence level and a size of the image input to the initial model.
30. The method of claim 29, wherein the model training method
35. The parameter display interface includes a confidence configuration icon, and the method further includes: and generating a trust configuration interface in response to a user's trigger operation on the trust configuration icon.
35. The method of claim 34, wherein the model training method
36. Training the initial model based on the training parameters to obtain a target model includes: determining, in response to a selection operation by a user on a reference model, the reference model corresponding to the selection operation as a target model; The reliability configuration interface includes a sample number corresponding to each defect and a selection icon corresponding to each defect, and the reliability configuration interface is used to configure a reliability of the defect corresponding to the selection operation in response to a user's decision operation on the selection icon.
36. The method of claim 35,
37. The method further comprises: generating and displaying a model training schedule, the schedule including a task cancel icon and a task detail icon; and stopping training of the initial model in response to a user's trigger operation on the task cancel icon.
30. The method of claim 29, wherein the model training method
38. The method further comprises: and generating and displaying a loss curve of the training process in response to a trigger operation by a user on the task details icon.
38. The method of claim 37, wherein the model training method
39. Training the initial model based on the training parameters to obtain a target model includes: obtaining a plurality of reference models based on the test policy, such that a target model is determined from the reference models based on a precision rate and a recall rate of each reference model; obtaining the precision rate and the recall rate of the plurality of reference models; and generating a reference model list.
30. The method of claim 29, wherein the model training method
40. Training the initial model based on the training parameters to obtain a target model includes: obtaining a plurality of the reference models based on the test policy, determining the confusion matrix of each of the reference models, and generating a display interface of the confusion matrix, so as to determine the target model from the reference models according to the confusion matrix; 35. The method of claim 34, wherein the model training method
41. The method further comprises: and updating the confidence levels based on the confusion matrix in response to a confidence level change operation by a user.
41. The method of claim 40, wherein the model training method
42. determining the target model from the reference model according to the confusion matrix, and determining, in response to a user's selection operation on the reference model, the reference model corresponding to the selection operation as a target model.
41. The method of claim 40, wherein the model training method
43. The total number of rounds of training is positively correlated to the number of samples.
30. The method of claim 29, wherein the model training method
44. Configuring a total number of rounds of training according to a predetermined rule based on the number of samples If the number of samples is less than or equal to 10,000, configuring the total number of rounds of training to be 300,000; If the number of samples is greater than 10,000, configuring the total number of rounds of the training with the following formula for determining the total number of rounds, wherein the formula for determining the total number of rounds is: Y=300000+INT(X / 10000)×b, Y denotes the total number of training rounds, X denotes the number of samples, and X is equal to or greater than 10,000. INT is a rounding function, b is a growth factor and is a fixed value, and b is equal to or greater than 30,000 and equal to or less than 70,000.
44. A method for training a model according to claim 29 or 43.
45. The number of rounds during the learning rate reduction is positively correlated with the total number of rounds in the training, and the number of rounds during the test is equal to or greater than the number of rounds during the first learning rate reduction and equal to or less than the total number of rounds in the training.
30. The method of claim 29, wherein the model training method
46. The learning rate is lowered multiple times, and at least two tests are performed within a predetermined range of the number of rounds of the second learning rate descent.
46. The method of claim 45,
47. The number of times the learning rate is decreased is three, and at least three tests are performed within a predetermined range of the number of rounds of the second learning rate decrease.
47. The method of claim 46, wherein the model training method
48. The learning rate reduction policy includes a learning rate reduction method and a learning rate reduction range.
30. The method of claim 29, wherein the model training method
49. 1. A defect detection system including a data management module, a training management module, and a model management module, the data management module is configured to store and manage sample data; The training management module is configured to execute the defect detection method according to any one of claims 1 to 21, the model training method according to any one of claims 21 to 28, or the model training method according to any one of claims 29 to 48, The model management module is configured to store, display, and manage the target models. A defect detection system comprising:
50. Includes a user management module, The user management module is configured to add, delete, change, view, manage authority, and / or manage passwords for user information.
50. The defect detection system of claim 49.
51. An electronic device including a processor and a memory that stores one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the defect detection method according to any one of claims 1 to 21, the model training method according to any one of claims 21 to 28, or the model training method according to any one of claims 29 to 48.
1. An electronic device comprising:
Citation Information
Patent Citations
An emulsion pump defect detection method based on deep learning
CN109559298A