Picture reinspection method, electronic equipment and storage medium
The image characteristics of the defective AOI detection equipment are extracted through the benchmark network and branch network framework, and the problem of low manual re-inspection efficiency of AOI detection equipment is solved, achieving efficient and accurate re-inspection effect.
Patent Information
- Application Number
- CN202410034990.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-09
- Publication Date
- 2025-07-11
AI Technical Summary
The existing AOI testing equipment needs manual re-inspection after detecting defective products, resulting in low labor waste and re-inspection efficiency, and the existing re-inspection model is low efficiency and pass rate.
The reference network and branch network framework are adopted to extract target image features through multiple encoders, determine the target branch network for rechecking, improve feature extraction efficiency and accuracy, and avoid misjudgment.
It improves the efficiency of re-inspection of pictures, reduces misjudgment, reduces labor costs, and ensures the accuracy of re-inspection.
Smart Images

Figure CN120293979A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing, and particularly to a method for re-inspecting pictures, an electronic device, and a storage medium. Background Art
[0002] An AOI (Automated Optical Inspection) device refers to an automated optical inspection device, which is a device for detecting common defects encountered in welding production based on optical principles. When the AOI inspection device is used to detect defects, it can automatically scan the PCB board through a camera to collect images, and then compare them with the qualified images pre-stored in the database, so as to check the defects on the PCB board, and display or mark the defects through a display or automatic identification for maintenance personnel to repair.
[0003] Currently, a detector is usually arranged beside the AOI inspection device at the factory end to be responsible for the operation of one AOI inspection device. When the AOI inspection device detects defective products, the detector manually transports the defective products to the re-inspection station, and finally the detector at the re-inspection station re-inspects the defective products. However, since there are many AOI inspection devices at the factory end, arranging a detector for each AOI inspection device will waste a lot of manpower, and it will also waste a lot of time for the detector to transport defective products each time and for the re-inspector to re-inspect the defective products. To solve such problems, an AOI re-judgment model is used in related technologies to re-judge defective products, but such methods require a large amount of data pre-processing during the re-judgment process, which affects the efficiency of the AOI re-judgment model on the one hand, and on the other hand, leads to too many over-killed pictures and a low pass rate. Summary of the Invention
[0004] Embodiments of this application disclose a method for re-inspecting pictures, an electronic device, and a storage medium, which solve the technical problem of low efficiency in re-inspecting pictures.
[0005] This application provides a method for re-inspecting pictures, and the method includes: receiving a target picture of a product that has been preliminarily screened by an automated optical inspection (AOI) device and whose preliminary screening result indicates unqualified; obtaining the picture detection type of the AOI device; inputting the target picture into a reference network to obtain an encoded vector of the target picture; determining a target branch network from multiple branch networks based on the picture detection type; and inputting the encoded vector into the target branch network to obtain a re-inspection result for the target picture.
[0006] In some embodiments of this application, the step of inputting the target picture into a reference network to obtain an encoded vector of the target picture includes: inputting the target picture into the reference network, and sequentially performing feature extraction through a first encoder, a second encoder, a third encoder, and a fourth encoder of the reference network to obtain an encoded vector of the target picture.
[0007] In some embodiments of the present application, inputting the target picture into the reference network and sequentially passing through the first encoder, the second encoder, the third encoder, and the fourth encoder of the reference network for feature extraction to obtain the encoded vector of the target picture includes: inputting the target picture into the first encoder to obtain a first global feature vector; inputting the first global feature vector into the second encoder to obtain a second global feature vector; inputting the second global feature vector into the third encoder to obtain a third global feature vector; and inputting the third global feature vector into the fourth encoder to obtain the encoded vector.
[0008] In some embodiments of the present application, inputting the target picture into the first encoder to obtain a first global feature vector includes: performing feature extraction on the target picture to obtain initial feature data; performing dimensionality reduction processing on the initial feature data to obtain dimensionality-reduced data; performing normalization processing on the dimensionality-reduced data to obtain first data; obtaining a first fusion feature of the first data according to the multi-head attention mechanism of the first encoder; calculating a first sum value of the first fusion feature and the dimensionality-reduced data, and performing normalization processing on the first sum value to obtain second data; performing a linear transformation on the second data by using the multi-layer perceptron model of the first encoder to obtain first transformed data; and calculating a second sum value of the first transformed data and the first sum value to obtain the first global feature vector.
[0009] In some embodiments of the present application, inputting the first global feature vector into the second encoder to obtain a second global feature vector includes: performing normalization processing on the first global feature vector to obtain third data; obtaining a second fusion feature of the third data according to the multi-head attention mechanism of the second encoder; calculating a third sum value of the second fusion feature and the first global feature vector, and performing normalization processing on the third sum value to obtain fourth data; performing a linear transformation on the fourth data by using the multi-layer perceptron model of the second encoder to obtain second transformed data; and calculating a fourth sum value of the second transformed data and the third sum value to obtain the second global feature vector.
[0010] In some embodiments of the present application, inputting the encoded vector into the target branch network to obtain a re-inspection result of the target picture includes: inputting the encoded vector into the first branch sub-network of the target branch network to obtain a first feature; inputting the first feature into the second branch sub-network of the target branch network to obtain a second feature; inputting the second feature into the third branch sub-network of the target branch network to obtain a third feature; performing dimensionality reduction on the third feature and performing a linear transformation by using a multi-layer perceptron model to obtain the re-inspection result.
[0011] In some embodiments of the present application, after obtaining the re-inspection result of the target picture, the method further includes: if the re-inspection result indicates that the product is unqualified, sending a first prompt message indicating that the product is unqualified to prompt the user to inspect the product corresponding to the target picture; if the re-inspection result indicates that the product is qualified, sending a second prompt message indicating that the product is qualified to prompt the user to transfer the product corresponding to the target picture to the product library.
[0012] In some embodiments of the present application, obtaining the picture detection type of the AOI device includes: determining the picture detection type according to the model of the AOI device.
[0013] The present application also provides an electronic device, which includes a processor and a memory. The processor is used to implement the picture re-inspection method when executing the computer program stored in the memory.
[0014] The present application also provides a computer-readable storage medium, on which a computer program is stored. The computer program is used to implement the picture re-inspection method when executed by a processor.
[0015] In the picture re-inspection method provided by the present application, when receiving a target picture of a product that has passed the initial screening by an automatic optical inspection (AOI) device and the initial screening result indicates unqualified, the picture detection type of the AOI device can be obtained to determine a target branch network from multiple branch networks. After determining the target picture, in order to extract more accurate features from the target picture, the target picture is input into a reference network for feature extraction to obtain an encoded vector of the target picture. Based on the determined target branch network, the encoded vector is input into the target branch network, thereby obtaining the re-inspection result of the target picture, which can improve the re-inspection efficiency of the picture and avoid misjudgment. Description of the Drawings
[0016] Figure 1 is a schematic diagram of the application scenario of the picture re-inspection method provided by the embodiments of the present application.
[0017] Figure 2 is a flowchart of the picture re-inspection method provided by the embodiments of the present application.
[0018] Figure 3 is a flowchart of the use of the reference network provided by the embodiments of the present application.
[0019] Figure 4 is a schematic structural diagram of the first encoder provided by the embodiments of the present application.
[0020] Figure 5 is a schematic structural diagram of the reference network provided by the embodiments of the present application.
[0021] Figure 6 It is a flowchart of the use of the branch network provided by the embodiments of the present application.
[0022] Figure 7 It is a schematic structural diagram of the first branch sub-network provided by the embodiments of the present application.
[0023] Figure 8 It is a schematic structural diagram of the branch network provided by the embodiments of the present application. Detailed implementation manners
[0024] For ease of understanding, some explanations of concepts related to the embodiments of the present application are exemplarily given for reference.
[0025] It should be noted that "at least one" in the present application means one or more, and "a plurality" means two or more than two. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B may be singular or plural. Terms such as "first", "second", "third", "fourth", etc. (if any) in the specification, claims and drawings of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0026] The AOI (Automated Optical Inspection) device refers to an automated optical inspection device, which is a device for detecting common defects encountered in welding production based on optical principles. When the AOI inspection device is used to detect defects, it can automatically scan the PCB board through a camera to collect images, and then compare them with the qualified images pre-stored in the database, so as to check the defects on the PCB board, and display or mark the defects through a display or automatic identification for repair personnel to repair.
[0027] Currently, a detector is usually arranged beside the AOI inspection device at the factory end to be responsible for the operation of one AOI inspection device. When the AOI inspection device detects defective products, the detector manually transports the defective products to the re-inspection station, and finally the detector at the re-inspection station reinspects the defective products. However, since there are many AOI inspection devices at the factory end, arranging a detector for each AOI inspection device will waste a lot of manpower, and the detector's transportation of defective products each time and the reinspection of defective products by the reinspector will also waste a lot of time. To solve such problems, an AOI rejudgment model is used in the related art to rejudge defective products, but such methods require a large amount of data preprocessing during the rejudgment process, which affects the efficiency of the AOI rejudgment model on the one hand, and on the other hand, will result in too many over-killed pictures, resulting in a low pass rate.
[0028] To better understand the image re-inspection method, electronic device, and storage medium provided by the embodiments of the present application, the application scenario of the image re-inspection method of the present application will be described first below.
[0029] Figure 1 It is a schematic diagram of the application scenario of the image re-inspection method provided by the embodiments of the present application. The image re-inspection method provided by the embodiments of the present application is applied to the electronic device 10. The electronic device 10 is communicatively connected to the detection device 20, which can be wired network communication or wireless network communication. Among them, the wired network can be any one of a local area network, a metropolitan area network, and a wide area network, and the wireless network can be any one of Bluetooth Technology, Wireless Fidelity (Wi-Fi), Near Field Communication (NFC), ZigBee Wireless Networks (ZigBee), Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Wireless Universal Serial Bus (USB), etc.
[0030] The electronic device 10 can be a computer device, a mobile phone, a tablet computer, a Personal Digital Assistant (PDA), etc. The electronic device 10 includes, but is not limited to, a memory 120 and at least one processor 130 connected by a communication bus 110.
[0031] The detection device 20 can be an Automated Optical Inspection (AOI) device, a device for detecting common defects in welding production. For example, whether the interfaces on the circuit board are accurate and whether there are connection errors.
[0032] In one example, the electronic device 10 is used to receive the target image sent by the detection device 20. The target image has been detected by the detection device 20, and the detection result (pre-screening result) output by the detection device 20 represents the target image of the unqualified product.
[0033] Figure 1 This is only an example of the electronic device 10 and does not limit the electronic device 10. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, the electronic device 10 may also include input / output devices, network access devices, etc.
[0034] To solve the above problems, please refer to Figure 2 as shownFigure 2 is a flowchart of the picture re-inspection method provided by an embodiment of this application, which is applied to an electronic device (such as Figure 1 the electronic device 10). According to different requirements, the order of steps in this flowchart can be changed, and some steps can be omitted.
[0035] Step S201: Receive the target picture of a product that has passed the initial screening by an Automatic Optical Inspection (AOI) device and whose initial screening result indicates non-conformance.
[0036] In some embodiments of this application, AOI devices are usually used to detect manufacturing defects of electronic products. For example, in the Surface Mount Technology (SMT) industry, AOI devices are used to detect the surface mounting quality and soldering quality of Surface Mounted Device (SMD) / Screw Jack (SJ) products. Through high-speed and high-precision vision processing technology, AOI devices automatically detect soldering defects and mounting errors on the Printed Circuit Board (PCB) to achieve quality control of SMD / SJ products and prevent defective products from flowing into subsequent processes. However, due to overly strict detection conditions, AOI devices may also determine defects within the error range as non-conforming. Therefore, to avoid misjudgment, this application uses an electronic device based on image processing to re-inspect products.
[0037] In some embodiments of this application, the electronic device can receive the target picture of a product that has been initially screened as non-conforming by the AOI device, and thus can further process the target picture without having to re-take pictures of the product, which can reduce the detection cost to a certain extent.
[0038] Step S202: Obtain the picture detection type of the AOI device.
[0039] In some embodiments of this application, each production machine can be configured with multiple AOI devices for quality inspection. The production machine can be a device for producing electronic products, such as a pick-and-place machine, a soldering device, etc. Electronic products can be circuit boards, mechanical parts, etc.
[0040] In some embodiments of this application, each AOI device can be set with different picture detection types for special detection of different defects. For example, the picture detection type can be soldering defects, pad defects, component placement defects, PCB board defects, etc. of electronic products. Among them, soldering defects can be bridging, cold soldering, dry soldering, etc.; pad defects can be insufficient tin on the pad, excessive tin on the pad, pad contamination, etc.; component placement defects can be wrong component, offset, reverse, etc.; PCB board defects can be scratches on the PCB board, warping of the PCB board, deformation of the PCB board, etc.
[0041] In one example, a first AOI device for detecting welding defects can be set beside the production machine, a second AOI device for detecting pad defects can be set, a third AOI device for detecting component placement defects can be set, and a fourth AOI device for detecting PCB board defects can be set.
[0042] Step S203: Input the target picture into the reference network to obtain the encoding vector of the target picture.
[0043] In some embodiments of the present application, the reference network can be composed of multiple encoders (Transformer Encoder) for extracting image features. To ensure the accuracy of feature extraction while reducing the time for feature extraction, multiple encoders can be used to form the reference network. For example, a reference network can be formed by 4 encoders, but the actual application is not limited to this.
[0044] In some embodiments of the present application, after obtaining the target picture, the electronic device can input the target picture into the reference network, and perform feature extraction through the first encoder, second encoder, third encoder, and fourth encoder of the reference network in sequence, so as to obtain the encoding vector of the target picture.
[0045] Step S204: Determine the target branch network from multiple branch networks based on the picture detection type.
[0046] In some embodiments of the present application, when the electronic device receives the target picture sent by the detection device, it can obtain the model of the detection device, and thus can determine the picture detection type according to the model of the detection device.
[0047] In one example, when the electronic device receives the target picture, it can obtain the Internet Protocol (IP) address of the detection device that sends the target picture. The electronic device can determine the model of the detection device according to the IP address. Among them, the IP address, corresponding model, and corresponding picture detection type of each detection device can be pre-stored in the electronic device, so that the corresponding model can be queried according to the IP address, and then the corresponding picture detection type can be obtained.
[0048] In another example, when the electronic device receives the target picture, it can also send an inquiry instruction to the detection device, and the inquiry instruction instructs the detection device to feedback the model of the detection device. When the electronic device receives the model feedback by the detection device, it can find the corresponding picture detection type from the pre-stored relationship library based on the model. Among them, the relationship library can be a database of the corresponding relationship between the model of the detection device and the picture detection type.
[0049] In some embodiments of the present application, after determining the picture detection type, the electronic device may select a corresponding branch network from multiple pre-stored branch networks as the target branch network. The branch network may be composed of multiple Convolutional Neural Network (CNN) structures to improve the prediction performance of the branch network. Among them, in the electronic device, a corresponding relationship may be established between the branch network and the picture detection type. When the electronic device receives the picture detection type, it may, based on the corresponding relationship, find the target branch network corresponding to the picture detection type, so that when the detection by the detection device fails, the branch network for detecting the same type of picture can be used for re-inspection. Each branch sub-network has a corresponding picture detection type for predicting whether the picture meets the requirements. The corresponding branch sub-network may be pre-trained based on the picture detection type, so that the branch sub-network can be used to detect the target picture of this picture detection type.
[0050] In some embodiments of the present application, the branch network may be composed of multiple CNN structures. For example, the first branch sub-network, the second branch sub-network, and the third branch sub-network are not limited in this application.
[0051] Step S205: Input the encoded vector into the target branch network to obtain the re-inspection result of the target picture.
[0052] In some embodiments of the present application, after determining the target branch network, the electronic device may input the encoded vector into the target branch network for re-inspection. After passing through the first branch sub-network, the second branch sub-network, and the third branch sub-network in the target branch network, the re-inspection result of the target picture is obtained.
[0053] In some embodiments of the present application, if the re-inspection result of the electronic device indicates that the product is unqualified, a first prompt message indicating that the product is unqualified may be sent to prompt the user to inspect the product corresponding to the target picture. For example, manual inspection or moving the unqualified product to the repair station is not limited in this application. If the re-inspection result of the electronic device indicates that the product is qualified, a second prompt message indicating that the product is qualified is sent to prompt the user to move the product corresponding to the target picture to the product library for further production or use.
[0054] In an embodiment of the present application, on the one hand, when receiving a target image of a product that has passed the initial screening by an Automatic Optical Inspection (AOI) device and the initial screening result indicates non-conformance, the image detection type of the AOI device can be obtained to determine a target branch network from multiple branch networks. After determining the target image, in order to extract more accurate features from the target image, the target image is input into a reference network for feature extraction to obtain an encoded vector of the target image. Based on the determined target branch network, the encoded vector is input into the target branch network, thereby obtaining a re-inspection result for the target image, which can improve the re-inspection efficiency of the image and avoid misjudgment. On the other hand, it is possible to further re-inspect the target image by constructing a framework of the reference network and the branch network. During the re-inspection process, there is no need to perform complex calculation processes on the data, which improves the re-inspection efficiency while ensuring the re-inspection accuracy.
[0055] Figure 3 is a flowchart of the use of the reference network provided by an embodiment of the present application. In order to improve the Figure 2 accuracy of the re-inspection of the embodiment shown, before inputting the target image into the target branch network, the input (encoded vector) input into the target branch network can be optimized so that the target branch network can output the re-inspection result more accurately, as shown in the Figure 3 embodiment shown, including the following steps:
[0056] Step S301, input the target image into a first encoder to obtain a first global feature vector.
[0057] Figure 4 is a schematic structural diagram of the first encoder provided by an embodiment of the present application.
[0058] In some embodiments of the present application, as shown in Figure 4 , the first encoder includes an input layer (EmbeddedPatches), a normalization layer (Layer Norm), an attention mechanism layer (Multi-Head Attention), and a multi-layer perceptron layer (Multi-Layer Perceptron, MLP). Among them, the input layer is used to extract initial feature data of the input target image, the normalization layer is used to accelerate the training and convergence speed of the reference network and prevent overfitting, the attention mechanism layer is used to extract information and feature fusion of the target image, and the multi-layer perceptron layer is used to transform the dimension of the features, which is convenient for stacking with other encoders of the reference network while extracting features.
[0059] In some embodiments of the present application, the target image is input into the input layer, and feature extraction is performed on the target image in the input layer to obtain initial feature data. The initial feature data is subjected to dimensionality reduction processing to obtain dimensionality-reduced data. For example, the initial feature data is 4-dimensional data and the dimensionality-reduced data is 3-dimensional data. The dimensionality-reduced data is input into the normalization layer, and the dimensionality-reduced data is normalized in the normalization layer to obtain the first data after normalization processing. The first data is input into the attention mechanism layer, and in the attention mechanism layer, the multi-head attention mechanism is used to perform feature fusion on the first data to obtain the first fusion feature. The first sum value of the first fusion feature and the dimensionality-reduced data is calculated, and the first sum value is input into the normalization layer to obtain the second data. The second data is input into the multi-layer perceptron layer, and the multi-layer perceptron model of the multi-layer perceptron layer is used to perform a linear transformation on the second data to obtain the first transformed data. The second sum value of the first transformed data and the first sum value is calculated, and the second sum value is used as the first global feature vector.
[0060] In some embodiments of the present application, the first global feature vector may include the basic features of the target image, such as color, shape, texture, etc.
[0061] Step S302: Input the first global feature vector into the second encoder to obtain the second global feature vector.
[0062] In some embodiments of the present application, the second encoder has the same structure as the first encoder. The first encoder can be used as the input layer of the second encoder. Then, the first global feature vector output by the first encoder is input into the second encoder, and the first global variable is input into the normalization layer. The first global feature vector is normalized in the normalization layer to obtain the third data. The third data is input into the attention mechanism layer, and in the attention mechanism layer, the multi-head attention mechanism is used to perform feature fusion on the third data to obtain the second fusion feature. The third sum value of the second fusion feature and the first global feature vector is calculated, and the third sum value is input into the normalization layer for normalization processing to obtain the fourth data. The fourth data is input into the multi-layer perceptron layer, and the multi-layer perceptron model of the multi-layer perceptron layer is used to perform a linear transformation on the fourth data to obtain the second transformed data. The fourth sum value of the second transformed data and the third sum value is calculated, and the fourth sum value is used as the second global feature vector.
[0063] In some embodiments of the present application, the second global feature vector may be an optimization of the first global feature vector, and the second global feature includes deeper features, such as contours, edges, etc.
[0064] Step S303: Input the second global feature vector into the third encoder to obtain the third global feature vector.
[0065] In some embodiments of the present application, the structure of the third encoder is the same as that of the second encoder. The second encoder can be used as the input layer of the third encoder. Then, the second global feature vector output by the second encoder is input into the third encoder, and the second global feature vector is input into the normalization layer. In the normalization layer, the second global feature vector is normalized to obtain the fifth data. The fifth data is input into the attention mechanism layer. In the attention mechanism layer, the multi-head attention mechanism is used to perform feature fusion on the fifth data to obtain the third fusion feature. The fifth sum value of the third fusion feature and the second global feature vector is calculated, and the fifth sum value is input into the normalization layer for normalization to obtain the sixth data. The sixth data is input into the multi-layer perceptron layer. The multi-layer perceptron model of the multi-layer perceptron layer is used to perform a linear transformation on the sixth data to obtain the third transformed data. The sixth sum value of the third transformed data and the fifth sum value is calculated, and the sixth sum value is used as the third global feature vector.
[0066] In some embodiments of the present application, the third global feature vector can be an optimization of the second global feature vector, and the third global feature includes more advanced features, such as categories.
[0067] Step S304: Input the third global feature vector into the fourth encoder to obtain an encoded vector.
[0068] In some embodiments of the present application, the structure of the fourth encoder is the same as that of the third encoder. The third encoder can be used as the input layer of the fourth encoder. Then, the third global feature vector output by the third encoder is input into the fourth encoder, and the third global variable is input into the normalization layer. In the normalization layer, the third global feature vector is normalized to obtain the seventh data. The seventh data is input into the attention mechanism layer. In the attention mechanism layer, the multi-head attention mechanism is used to perform feature fusion on the seventh data to obtain the fourth fusion feature. The seventh sum value of the fourth fusion feature and the third global feature vector is calculated, and the seventh sum value is input into the normalization layer for normalization to obtain the eighth data. The eighth data is input into the multi-layer perceptron layer. The multi-layer perceptron model of the multi-layer perceptron layer is used to perform a linear transformation on the eighth data to obtain the fourth transformed data. The eighth sum value of the fourth transformed data and the seventh sum value is calculated, and the eighth sum value is used as the encoded vector, where the dimension of the encoded vector is 3D. The encoded vector can include the features that have the greatest impact on the classification result, providing support for subsequent classification (such as Figure 5 the embodiments shown).
[0069] Figure 5 is a schematic structural diagram of the benchmark network provided by the embodiments of the present application. As Figure 5As shown, the structures of the first encoder, the second encoder, the third encoder, and the fourth encoder are the same. Among them, the first encoder can serve as the input layer of the second encoder, the second encoder can serve as the input layer of the third encoder, and the third encoder can serve as the input layer of the fourth encoder.
[0070] In the embodiments of the present application, in practical applications, the more the number of encoders, the longer the parameters to be trained and the training time, and the longer the time to obtain the encoded vector. Therefore, in the embodiments of the present application, the benchmark network is composed of 4 encoders, which can balance accuracy and efficiency. In addition, by continuously iteratively extracting features, the encoded vector output by the benchmark network includes the features that have the greatest impact on the classification result, which can improve the accuracy of subsequent classification to a certain extent.
[0071] Figure 6 is the usage flowchart of the branch network provided by the embodiments of the present application. As shown in Figure 3 After determining the encoded vector in the shown embodiment, the encoded vector is input into the target branch network for processing. As shown in Figure 6 shown, it includes the following steps:
[0072] Step S601, input the encoded vector into the first branch sub-network of the target branch network to obtain the first feature.
[0073] Figure 7 is the structural schematic diagram of the first branch sub-network provided by the embodiments of the present application.
[0074] In some embodiments of the present application, after the benchmark network outputs the encoded vector, the encoded vector can be used as the input vector of the target branch network. The encoded vector is input into the first branch sub-network of the target branch network for processing. Among them, the first branch sub-network is as shown in Figure 7 shown. The first branch sub-network includes an input layer, a convolutional layer (Conv), a pooling layer (MaxPooling), an activation layer (ReLu), and a batch normalization layer (Batch Norm). Among them, the convolutional layer is used to perform convolution on the encoded vector to obtain more features. The pooling layer is used to perform downsampling on the features, reduce the operation time, and perform dimensionality reduction processing on the features. The activation layer is used to convert the linear transformation into a non-linear transformation. The batch normalization layer is used to perform standardization processing on the features to solve the problem of numerical instability.
[0075] In some embodiments of the present application, since the encoded vector output by the reference network is 3-dimensional, the 3-dimensional encoded vector can be converted into a 4-dimensional encoded vector and then input into the convolutional layer of the first branch sub-network to obtain a first convolutional result. The first convolutional result is input into a pooling layer to obtain a first pooling result. The first pooling result is input into an activation layer to obtain a first activation result. The first activation result is input into a batch normalization layer to obtain a first feature. The first feature can be the primary features (such as spatial hierarchical structure information and features such as color and texture) extracted from the encoded vector, which are used to distinguish different objects or scenes.
[0076] Step S602: Input the first feature into the second branch sub-network of the target branch network to obtain a second feature.
[0077] In some embodiments of the present application, the structure of the second branch sub-network is the same as that of the first branch sub-network. The output of the first branch sub-network is used as the input of the second branch sub-network. Then, the first feature is input into the convolutional layer of the second branch sub-network to obtain a second convolutional result. The second convolutional result is input into a pooling layer to obtain a second pooling result. The second pooling result is input into an activation layer to obtain a second activation result. The second activation result is input into a batch normalization layer to obtain a second feature. The second feature can be more complex features further extracted from the first feature, and this process will gradually abstract the high-level information in the image.
[0078] Step S603: Input the second feature into the third branch sub-network of the target branch network to obtain a third feature.
[0079] In some embodiments of the present application, the structure of the third branch sub-network is the same as that of the second branch sub-network. The output of the second branch sub-network is used as the input of the third branch sub-network. Then, the second feature is input into the convolutional layer of the third branch sub-network to obtain a third convolutional result. The third convolutional result is input into a pooling layer to obtain a third pooling result. The third pooling result is input into an activation layer to obtain a third activation result. The third activation result is input into a batch normalization layer to obtain a third feature. The third feature can be more accurate classification features further extracted from the second feature, which are used to correctly identify the category.
[0080] Step S604: Perform dimensionality reduction on the third feature and perform a linear transformation using a multi-layer perceptron model to obtain a re-inspection result.
[0081] In some embodiments of the present application, the dimension of the third feature is 4-dimensional. The 4-dimensional third feature can be reduced to 2-dimensional, and a multi-layer perceptron model is used to perform a linear transformation on the 2-dimensional third feature to obtain a re-inspection result. The re-inspection result includes the probability of the product being qualified and the probability of the product being unqualified. For example, if the probability of the product being qualified is 0.9, then the probability of the product being unqualified is 0.1. At this time, the re-inspection result indicates that the product is qualified.
[0082] In an embodiment of the present application, by setting multiple branch sub-networks in the branch network, the accuracy of recognition can be improved. In an embodiment of the present application, three branch sub-networks are used to process the encoded vector output by the benchmark network, which can balance the efficiency and accuracy of prediction and reduce the false detection rate to a certain extent. In addition, the use of multiple branch sub-networks in the present application helps to improve the accuracy and robustness of classification. Stacking multiple branch sub-networks (such as CNN Module) can also increase the depth of the branch network, thereby enhancing the learning ability and representation ability of the branch network.
[0083] Figure 8 is a schematic structural diagram of the branch network provided by an embodiment of the present application. As Figure 8 shown, it includes a first branch sub-network, a second branch sub-network, and a third branch sub-network. Among them, the first branch sub-network, the second branch sub-network, and the third branch sub-network all include a convolutional layer, a pooling layer, an activation layer, and a batch normalization layer. The first branch sub-network serves as the input layer of the second branch sub-network, and the second branch sub-network serves as the input layer of the third branch sub-network.
[0084] Please continue to refer to Figure 1 , in this embodiment, the memory 120 may be an internal memory of the electronic device 10, that is, a memory built into the electronic device 10. In other embodiments, the memory 120 may also be an external memory of the electronic device 10, that is, a memory externally connected to the electronic device 10.
[0085] In some embodiments, the memory 120 is used to store program codes and various data, and realizes the high-speed and automatic access of programs or data during the operation of the electronic device 10.
[0086] The memory 120 may include a random access memory, and may also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0087] In one embodiment, the processor 130 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any other conventional processor, etc.
[0088] If the program code and various data in the memory 120 are implemented in the form of software functional units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-described embodiment methods of the present application, for example, the picture recheck method, it may also be completed by instructing relevant hardware through a computer program. The computer program may be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above-described method embodiments may be implemented. Among them, the computer program includes computer program code, and the computer program code may be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), etc.
[0089] It can be understood that the above-described module division is a logical function division, and there may be other division methods in actual implementation. In addition, in each embodiment of the present application, the functional modules may be integrated in the same processing unit, or each module may exist physically alone, or two or more modules may be integrated in the same unit. The above-mentioned integrated modules may be implemented in the form of hardware, or in the form of a combination of hardware and software functional modules.
[0090] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and not to limit them. Although the present application has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present application may be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present application.
Claims
1. A method for rechecking pictures, characterized in that, The method includes: Receiving a target image of a product that has passed the initial screening by an automatic optical inspection (AOI) device and whose initial screening result indicates non - compliance; Obtaining the image detection type of the AOI device; Inputting the target image into a reference network to obtain an encoded vector of the target image; Based on the image detection type, determining a target branch network from multiple branch networks; Inputting the encoded vector into the target branch network to obtain a re - inspection result for the target image.
2. The picture rechecking method according to claim 1, wherein The step of inputting the target image into the reference network to obtain an encoded vector of the target image includes: Inputting the target image into the reference network, and successively performing feature extraction through the first encoder, second encoder, third encoder, and fourth encoder of the reference network to obtain an encoded vector of the target image.
3. The picture rechecking method according to claim 2, wherein The step of inputting the target image into the reference network, and successively performing feature extraction through the first encoder, second encoder, third encoder, and fourth encoder of the reference network to obtain an encoded vector of the target image includes: Inputting the target image into the first encoder to obtain a first global feature vector; Inputting the first global feature vector into the second encoder to obtain a second global feature vector; Inputting the second global feature vector into the third encoder to obtain a third global feature vector; Inputting the third global feature vector into the fourth encoder to obtain the encoded vector.
4. The picture rechecking method according to claim 3, characterized in that, The step of inputting the target image into the first encoder to obtain a first global feature vector includes: Performing feature extraction on the target image to obtain initial feature data; Performing dimensionality reduction processing on the initial feature data to obtain dimensionality - reduced data; Performing normalization processing on the dimensionality - reduced data to obtain first data; According to the multi - head attention mechanism of the first encoder, obtaining a first fusion feature of the first data; Calculating a first sum value of the first fusion feature and the dimensionality - reduced data, and performing normalization processing on the first sum value to obtain second data; Using the multi - layer perceptron model of the first encoder to perform a linear transformation on the second data to obtain first transformed data; Calculating a second sum value of the first transformed data and the first sum value to obtain the first global feature vector.
5. The picture rechecking method according to claim 4, wherein, The step of inputting the first global feature vector into the second encoder to obtain a second global feature vector includes: Performing normalization processing on the first global feature vector to obtain third data; According to the multi - head attention mechanism of the second encoder, obtaining a second fusion feature of the third data; Calculating a third sum value of the second fusion feature and the first global feature vector, and performing normalization processing on the third sum value to obtain fourth data; Using the multi - layer perceptron model of the second encoder to perform a linear transformation on the fourth data to obtain second transformed data; Calculating a fourth sum value of the second transformed data and the third sum value to obtain the second global feature vector.
6. The picture rechecking method according to claim 1, wherein The step of inputting the encoded vector into the target branch network to obtain a re - inspection result for the target image includes: Inputting the encoded vector into the first branch sub - network of the target branch network to obtain a first feature; Input the first feature into the second branch sub-network of the target branch network to obtain a second feature; Input the second feature into the third branch sub-network of the target branch network to obtain a third feature; Perform dimensionality reduction on the third feature and perform a linear transformation using a multi-layer perceptron model to obtain the re-inspection result.
7. The picture rechecking method according to claim 1, wherein After obtaining the re-inspection result of the target picture, the method further includes: If the re-inspection result indicates that the product is unqualified, send a first prompt message indicating that the product is unqualified to prompt the user to inspect the product corresponding to the target picture; If the re-inspection result indicates that the product is qualified, send a second prompt message indicating that the product is qualified to prompt the user to transfer the product corresponding to the target picture to the product library.
8. The picture rechecking method according to claim 1, characterized in that The obtaining of the picture detection type of the AOI device includes: Determine the picture detection type according to the model of the AOI device.
9. An electronic device, characterized in that, The electronic device includes a processor and a memory, and the processor is configured to execute a computer program stored in the memory to implement the picture re-inspection method according to any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one instruction, and when the at least one instruction is executed by a processor, the picture re-inspection method according to any one of claims 1 to 8 is implemented.