An image data processing method, a computer device, and a readable storage medium

Through multi-view decision analysis and image data processing methods, combined with instance segmentation and sub-classification models, the accuracy and efficiency problems existing in manual detection are solved, and efficient and accurate quality detection of industrial components is achieved.

CN114331949BActive Publication Date: 2025-07-22TENCENT TECH SHANGHAI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111153279.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-29
Publication Date
2025-07-22
Estimated Expiration
2041-09-29

AI Technical Summary

Technical Problem

In the prior art, the quality inspection of industrial components relies on manual naked eye review, resulting in complex work, high inconsistency, slow speed and low detection accuracy, making it difficult to fully identify defects of components.

Method used

The image data processing method is adopted to identify and classify defect labeling areas of target images from multiple angles through multi-view decision analysis, combined with instance segmentation model and sub-categorization model, so as to realize multi-view decision-making and improve detection accuracy.

Benefits of technology

It improves the accuracy and efficiency of quality inspection, reduces defect missed inspection and misjudgment, and achieves efficient and accurate quality inspection of components.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114331949B_ABST
    Figure CN114331949B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide an image data processing method, a computer device, and a readable storage medium. The method relates to fields such as artificial intelligence, intelligent transportation, and assisted driving. The method includes: obtaining S defect annotation regions associated with N target images, and first defect output results respectively corresponding to the S defect annotation regions; according to the defect annotation region of target image L i and the image attribute information of target image L i , determining a second defect output result corresponding to the defect annotation region of target image L i ; based on the first defect output results respectively corresponding to the S defect annotation regions and the second defect output results respectively corresponding to the S defect annotation regions, performing multi-view decision analysis on the target object to obtain an object detection result of the target object. By using the present application, quality detection of the target object can be realized, and thus the accuracy of quality detection can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to an image data processing method, a computer device, and a readable storage medium. Background Art

[0002] For target objects (i.e., components) in industry, the existing defect quality inspection process is to use the combination of manual visual inspection and microscope review to perform quality inspection on the target images of components. It can be understood that manual quality inspection work is cumbersome and boring, which is likely to cause personnel loss; manual quality inspection is subjective and has obvious inconsistencies; manual quality inspection is slow, and the production efficiency is low. In addition, the surface structure of components is very complex, and it is inevitable that the manual quality inspection method will miss the defects in a certain position of the components, thereby reducing the accuracy of quality inspection. Summary of the Invention

[0003] The embodiments of this application provide an image data processing method, a computer device, and a readable storage medium, which can improve the accuracy of quality inspection.

[0004] On the one hand, the embodiments of this application provide an image data processing method, including:

[0005] Obtain S defect annotation regions associated with N target images, and first defect output results respectively corresponding to the S defect annotation regions; the N target images are obtained by N shooting components respectively shooting the same target object; the visual angles of the N target images are different from each other; N is a positive integer; S is a positive integer; the N target images include target image Li, where i is a positive integer less than or equal to N; i , i is a positive integer less than or equal to N;

[0006] According to the defect annotation region of target image Li i and the image attribute information of target image Li i , determine the second defect output result corresponding to the defect annotation region of target image Li i ;

[0007] Based on the first defect output results respectively corresponding to the S defect annotation regions and the second defect output results respectively corresponding to the S defect annotation regions, perform multi-view decision analysis on the target object to obtain the object detection result of the target object.

[0008] On the one hand, the embodiments of this application provide an image data processing device, including:

[0009] The first output module is used to obtain S defect annotation regions associated with N target images, and first defect output results respectively corresponding to the S defect annotation regions; the N target images are obtained by N shooting components respectively shooting the same target object; the visual angles of the N target images are different from each other; N is a positive integer; S is a positive integer; the N target images include target image L i , where i is a positive integer less than or equal to N;

[0010] The second output module is used to determine a second defect output result corresponding to the defect annotation region of target image L i according to the defect annotation region of target image L i and the image attribute information of target image L i ;

[0011] The decision analysis module is used to perform multi-view decision analysis on the target object based on the first defect output results respectively corresponding to the S defect annotation regions and the second defect output results respectively corresponding to the S defect annotation regions, so as to obtain an object detection result of the target object.

[0012] Among them, the first output module includes:

[0013] The image acquisition unit is used to acquire N target images associated with the target object and input the N target images into the instance segmentation model respectively;

[0014] The instance segmentation unit is used to perform instance segmentation on the N target images through the instance segmentation model to obtain S defect annotation regions associated with the N target images and first defect output results respectively corresponding to the S defect annotation regions.

[0015] Among them, the instance segmentation model includes a feature extraction sub-network, a region prediction sub-network and a defect recognition sub-network; the S defect annotation regions include M defect annotation regions in target image L i ; M is a positive integer less than or equal to S;

[0016] The instance segmentation unit includes:

[0017] The feature extraction sub-unit is used to input target image L i into the feature extraction sub-network, and perform feature extraction on target image L i through the feature extraction sub-network to obtain multi-resolution features corresponding to target image L i ;

[0018] The region prediction sub-unit is used to input the multi-resolution features corresponding to target image L i into the region prediction sub-network, and perform region prediction on target image L iPerform region prediction on the corresponding multi-resolution features to obtain the target image L i among the M regions of objects to be predicted;

[0019] A defect recognition subunit, configured to input the M regions of objects to be predicted and the multi-resolution features corresponding to the target image L i into a defect recognition sub-network, and perform defect recognition on the M regions of objects to be predicted and the multi-resolution features corresponding to the target image L i to obtain the instance segmentation results respectively corresponding to the M defect annotation regions, the first classification probabilities respectively corresponding to the M defect annotation regions, and the first classification information respectively corresponding to the M defect annotation regions;

[0020] A defect recognition subunit, configured to use the instance segmentation results respectively corresponding to the M defect annotation regions, the first classification probabilities respectively corresponding to the M defect annotation regions, and the first classification information respectively corresponding to the M defect annotation regions as the first defect output results respectively corresponding to the M defect annotation regions.

[0021] Among them, the defect recognition subunit is specifically configured to map the M regions of objects to be predicted to the multi-resolution features corresponding to the target image L i through the defect recognition sub-network to obtain the candidate region features respectively corresponding to the M regions of objects to be predicted;

[0022] A defect recognition subunit, specifically configured to perform feature alignment on the M candidate region features to obtain the aligned region features respectively corresponding to the M candidate region features;

[0023] A defect recognition subunit, specifically configured to perform a convolution operation on the M aligned region features to obtain the classification region features respectively corresponding to the M aligned region features and the segmentation region features respectively corresponding to the M aligned region features;

[0024] A defect recognition subunit, specifically configured to perform a fully connected operation on the M classification region features to determine the region features respectively corresponding to the M aligned region features and the classification features respectively corresponding to the M aligned region features, determine the M defect annotation regions based on the M region features, and determine the first classification probabilities respectively corresponding to the M defect annotation regions and the first classification information respectively corresponding to the M defect annotation regions based on the M classification features;

[0025] A defect recognition subunit, specifically configured to perform a convolution operation on the M segmentation region features to determine the segmentation features respectively corresponding to the M aligned region features, and determine the instance segmentation results respectively corresponding to the M defect annotation regions based on the M segmentation features.

[0026] Among them, the image attribute information of the target image L i includes the target image L iThe image serial number and the target image L i The corresponding image output feature;

[0027] The second output module includes:

[0028] The first determination unit is used to determine, according to the defect annotation area of the target image L i and the image serial number of the target image L i the defect output feature corresponding to the defect annotation area of the target image L i ;

[0029] The second determination unit is used to determine, according to the defect output feature corresponding to the defect annotation area of the target image L i and the image output feature corresponding to the target image L i the second defect output result corresponding to the defect annotation area of the target image L i .

[0030] Among them, the first determination unit includes:

[0031] The first determination subunit is used to determine the regional coordinates of the defect annotation area of the target image L i , generate, according to the regional coordinates and the image serial number of the target image L i the defect input feature corresponding to the defect annotation area of the target image L i , and input the defect input feature into the fine classification model; the fine classification model includes a perceptron subnetwork;

[0032] The second determination subunit is used to perform a full connection operation on the defect input feature through the perceptron subnetwork to determine the defect output feature corresponding to the defect annotation area of the target image L i .

[0033] Among them, the fine classification model further includes a feature recognition subnetwork;

[0034] The second determination unit includes:

[0035] The feature recognition subunit is used to input the target image L i into the feature recognition subnetwork, perform feature recognition on the target image L i through the feature recognition subnetwork, and obtain the image output feature corresponding to the target image L i ;

[0036] The feature fusion subunit is used to perform feature fusion on the defect output feature corresponding to the defect annotation area of the target image L i and the image output feature corresponding to the target image L i to obtain the fusion output feature corresponding to the defect annotation area of the target image L i ;

[0037] A region classification subunit, configured to determine a second defect output result corresponding to a defect annotation region of a target image L i based on a fusion output feature corresponding to the defect annotation region of the target image L i and a classifier of a fine classification model.

[0038] Specifically, the region classification subunit is configured to input a fusion output feature corresponding to a defect annotation region of the target image L i into the classifier of the fine classification model, and determine a matching degree between the fusion output feature corresponding to the defect annotation region of the target image L i and a sample output feature in the classifier; the matching degree is used to describe the probability that the defect annotation region of the target image L i belongs to a sample classification label corresponding to the sample output feature;

[0039] Specifically, the region classification subunit is configured to use the sample classification label corresponding to the sample output feature with the maximum matching degree as the second classification information corresponding to the defect annotation region of the target image L i and use the maximum matching degree as the second classification probability corresponding to the defect annotation region of the target image L i ;

[0040] Specifically, the region classification subunit is configured to use the second classification information corresponding to the defect annotation region of the target image L i and the second classification probability corresponding to the defect annotation region of the target image L i as the second defect output result corresponding to the defect annotation region of the target image L i .

[0041] Wherein, the decision analysis module includes:

[0042] A decision tree generation unit, configured to obtain business knowledge for multi-perspective decision analysis of a target object and target decision hyperparameters associated with the business knowledge, and generate a decision tree according to the business knowledge and the target decision hyperparameters;

[0043] A decision analysis unit, configured to perform multi-perspective decision analysis on N target images in a decision analysis model based on first defect output results respectively corresponding to S defect annotation regions, second defect output results respectively corresponding to the S defect annotation regions, and the decision tree, and obtain image detection results of the N target images respectively;

[0044] A result determination unit, configured to determine an object detection result of the target object according to the image detection results of the N target images respectively.

[0045] Wherein, the decision tree generation unit includes:

[0046] A set generation subunit for obtaining business knowledge and a hyperparameter search model for multi-perspective decision analysis of a target object, and generating a set of hyperparameters associated with the business knowledge through the hyperparameter search model; the set of hyperparameters includes one or more sets of decision hyperparameters; each set of decision hyperparameters in the one or more sets of decision hyperparameters includes one or more hyperparameters; the one or more sets of decision hyperparameters are used to balance at least two evaluation indicators corresponding to the decision analysis model.

[0047] A decision tree generation subunit for obtaining target decision hyperparameters that meet the hyperparameter acquisition conditions from the set of hyperparameters, and generating a decision tree according to the business knowledge and the target decision hyperparameters.

[0048] Among them, the target decision hyperparameters include instance segmentation hyperparameters, segmentation area hyperparameters, and fine classification hyperparameters.

[0049] The decision analysis unit includes:

[0050] A parameter acquisition unit for obtaining instance segmentation results corresponding to S defect annotation regions respectively from the first defect output results corresponding to the S defect annotation regions, and determining the defect region areas corresponding to the S defect annotation regions respectively according to the instance segmentation results corresponding to the S defect annotation regions.

[0051] A parameter acquisition unit for obtaining the first classification probabilities corresponding to the S defect annotation regions respectively, and the first classification information corresponding to the S defect annotation regions respectively, from the first defect output results corresponding to the S defect annotation regions, and obtaining the second classification probabilities corresponding to the S defect annotation regions respectively, and the second classification information corresponding to the S defect annotation regions respectively, from the second defect output results corresponding to the S defect annotation regions.

[0052] A decision analysis subunit for performing multi-perspective decision analysis on N target images in the decision analysis model according to the first classification information corresponding to the S defect annotation regions respectively, the second classification information corresponding to the S defect annotation regions respectively, the first classification probabilities corresponding to the S defect annotation regions respectively, the second classification probabilities corresponding to the S defect annotation regions respectively, the defect region areas corresponding to the S defect annotation regions respectively, the S defect annotation regions, and the instance segmentation hyperparameters, segmentation area hyperparameters, and fine classification hyperparameters indicated by the decision tree, to obtain the image detection results of the N target images respectively.

[0053] Among them, the first output module further includes:

[0054] A label acquisition unit for acquiring defect sample annotation regions, defect sample classification information, and sample boundary regions associated with the defect sample images.

[0055] A model output unit, configured to determine, in an initial instance segmentation model, a predicted defect annotation region associated with a defect sample image and a first predicted output result corresponding to the predicted defect annotation region;

[0056] A model training unit, configured to determine an instance segmentation loss value of the initial instance segmentation model according to a defect sample annotation region, defect sample classification information, a sample boundary region, the predicted defect annotation region, and the first predicted output result;

[0057] A model training unit, configured to adjust model parameters in the initial instance segmentation model according to the instance segmentation loss value, and when the adjusted initial instance segmentation model meets the model convergence condition, determine the adjusted initial instance segmentation model as the instance segmentation model.

[0058] Wherein, the first determination unit further includes:

[0059] A label acquisition subunit, configured to acquire a defect sample annotation region and defect sample classification information associated with the defect sample image, and acquire a normal sample annotation region and normal sample classification information associated with a normal sample image;

[0060] A model output subunit, configured to determine, in an initial fine classification model, a second predicted output result corresponding to the defect sample annotation region according to the defect sample annotation region and image attribute information of the defect sample image, and determine a first classification loss value of the initial fine classification model according to the second predicted output result corresponding to the defect sample annotation region and the defect sample classification information;

[0061] A model output subunit, configured to determine a second predicted output result corresponding to the normal sample annotation region according to the normal sample annotation region and image attribute information of the normal sample image, and determine a second classification loss value of the initial fine classification model according to the second predicted output result corresponding to the normal sample annotation region and the normal sample classification information;

[0062] A model training subunit, configured to determine a fine classification loss value of the initial fine classification model according to the first classification loss value and the second classification loss value;

[0063] A model training subunit, configured to adjust model parameters in the initial fine classification model according to the fine classification loss value, and when the adjusted initial fine classification model meets the model convergence condition, determine the adjusted initial fine classification model as the fine classification model.

[0064] An embodiment of the present application provides a computer device on the one hand, including: a processor and a memory;

[0065] The processor is connected to the memory. The memory is used to store a computer program. When the computer program is executed by the processor, the computer device executes the method provided by the embodiments of the present application.

[0066] On the one hand, the embodiments of the present application provide a computer-readable storage medium storing a computer program, which is adapted to be loaded and executed by a processor so that a computer device having the processor executes the method provided by the embodiments of the present application.

[0067] On the one hand, the embodiments of the present application provide a computer program product or a computer program. The computer program product or the computer program includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions so that the computer device executes the method provided by the embodiments of the present application.

[0068] In the embodiments of the present application, the computer device may obtain S defect annotation regions associated with N target images, and first defect output results respectively corresponding to the S defect annotation regions. Among them, the N target images are obtained by N shooting components respectively shooting the same target object, and the visual angles of the N target images are different from each other; both N and S here may be positive integers; the N target images include target image Li i , where i may be a positive integer less than or equal to N. Further, the computer device may determine target image Li i according to the defect annotation region of target image Li i and the image attribute information of target image Li iThe second defect output result corresponding to the defect annotation area. Further, the computer device can perform multi-perspective decision analysis on the target object based on the first defect output results corresponding to the S defect annotation areas and the second defect output results corresponding to the S defect annotation areas, to obtain the object detection result of the target object. Thus, it can be seen that the embodiments of the present application can perform rough quality detection on N target images associated with the target object, highly detect all defect annotation areas (i.e., S defect annotation areas) in the N target images, and then perform fine quality detection on the S defect annotation areas to further identify the S defect annotation areas. It can be understood that based on the first defect output result obtained from the rough quality detection and the second defect output result obtained from the fine quality detection, the object detection result of the decision target object can be inferred. Therefore, the embodiments of the present application can, through rough quality detection and fine quality detection, identify relatively accurate defect annotation areas and the defect output results corresponding to the defect annotation areas in the N target images, and then, through these accurately identified defect output results, achieve the quality detection of the target object targeted by the N target images, thereby improving the accuracy of quality detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0070] Figure 1 is a schematic structural diagram of a network architecture provided by an embodiment of the present application;

[0071] Figure 2 is a schematic diagram of a scenario for data interaction provided by an embodiment of the present application;

[0072] Figure 3 is a schematic flowchart of an image data processing method provided by an embodiment of the present application;

[0073] Figure 4 is a schematic structural diagram of a point design provided by an embodiment of the present application;

[0074] Figure 5a is a schematic architecture diagram of a defect quality inspection solution provided by an embodiment of the present application;

[0075] Figure 5b is a schematic architecture diagram of a defect quality inspection solution provided by an embodiment of the present application;

[0076] Figure 6It is a schematic structural diagram of a defect quality inspection solution provided by an embodiment of the present application;

[0077] Figure 7 It is a schematic flowchart of an image data processing method provided by an embodiment of the present application;

[0078] Figure 8 It is a schematic structural diagram of an instance segmentation model provided by an embodiment of the present application;

[0079] Figure 9 It is a schematic flowchart of an image data processing method provided by an embodiment of the present application;

[0080] Figure 10 It is a schematic structural diagram of a fine-grained classification model provided by an embodiment of the present application;

[0081] Figure 11 It is a schematic flowchart of an image data processing method provided by an embodiment of the present application;

[0082] Figure 12 It is a schematic flowchart of a process for generating a set of hyperparameters provided by an embodiment of the present application;

[0083] Figure 13 It is a schematic diagram of a scenario for performing defect quality inspection provided by an embodiment of the present application;

[0084] Figure 14 It is a schematic diagram of a scenario for comparing multiple models provided by an embodiment of the present application;

[0085] Figure 15 It is a schematic structural diagram of an image data processing device provided by an embodiment of the present application;

[0086] Figure 16 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0087] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0088] It should be understood that artificial intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.

[0089] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, autonomous driving, and intelligent transportation.

[0090] Among them, the solution provided by the embodiments of the present application mainly relates to the computer vision (CV) technology and machine learning (ML) technology of artificial intelligence.

[0091] Among them, computer vision is a science that studies how to make machines "see". Further, it refers to using cameras and computers to replace human eyes for object recognition, measurement, and other machine vision, and further performing graphic processing to make the computer process into an image that is more suitable for human eyes to observe or transmit to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to establish an artificial intelligence system that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, autonomous driving, intelligent transportation, and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.

[0092] Among them, Machine Learning is an interdisciplinary subject involving multiple fields, such as probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory, etc. It specializes in studying how computers simulate or implement human learning behaviors to acquire new knowledge or skills, and reorganize the existing knowledge structure to continuously improve its own performance. Machine Learning is the core of Artificial Intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of Artificial Intelligence. Machine Learning and Deep Learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. Among them, Deep Learning technology is a technology that uses deep neural network systems for machine learning.

[0093] Specifically, please refer to Figure 1 , Figure 1 which is a schematic structural diagram of a network architecture provided by an embodiment of the present application. As Figure 1 shown, as Figure 1 shown, the network architecture may include a server 2000 and a cluster of terminal devices. Among them, the cluster of terminal devices may specifically include one or more terminal devices, and the number of terminal devices in the cluster of terminal devices will not be limited here. As Figure 1 shown, the multiple terminal devices may specifically include terminal device 3000a, terminal device 3000b, terminal device 3000c,..., terminal device 3000n; terminal device 3000a, terminal device 3000b, terminal device 3000c,..., terminal device 3000n may be directly or indirectly network-connected to the server 2000 through wired or wireless communication methods, so that each terminal device can perform data interaction with the server 2000 through this network connection. In addition, terminal device 3000a, terminal device 3000b, terminal device 3000c,..., terminal device 3000n may be directly or indirectly network-connected to each other through wired or wireless communication methods, so that each terminal device can perform data interaction through this network connection.

[0094] Among them, each terminal device in the cluster of terminal devices may include: intelligent terminals with image data processing functions such as smart phones, tablet computers, notebook computers, desktop computers, smart home appliances, wearable devices, vehicle-mounted terminals, intelligent voice interaction devices, cameras, etc. For ease of understanding, an embodiment of the present application may select one or more terminal devices from the Figure 1 multiple terminal devices shown as the target terminal devices. For example, an embodiment of the present application may use Figure 1 the terminal device 3000a and the terminal device 3000c shown as the target terminal devices.

[0095] Among them, the server 2000 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.

[0096] Among them, a shooting component for collecting a target image associated with the target object can be integrally installed on the target terminal device. Here, the shooting component can be a camera for taking photos on the target terminal device. Among them, multiple cameras can be integrally installed on one target terminal device. In this embodiment of the application, multiple cameras on one target terminal device are taken as an example of a shooting component for illustration.

[0097] It can be understood that the target image can include a defective image and a non-defective image (i.e., a normal image). Here, the defective image can be an image of a component (i.e., the target object) taken by the camera with an NG (No Good) defect, and the non-defective image can be an image of a component taken by the camera with an OK (good) defect, or an image without a defect (i.e., OK defect and NG defect). Among them, the defect types (i.e., NG defect and OK defect) corresponding to the defects in the target image can be multiple. In this embodiment of the application, the defect types corresponding to the defects in the target image are not limited, and the number of defects in the target image is not limited in this embodiment of the application.

[0098] Among them, an OK defect is a defect that has no impact on the use of the product (i.e., the component), or a defect that can be eliminated through subsequent processing. For example, dirt and bright marks, etc.; an NG defect is a defect that affects the use of the product functionally. For example, cracks, material shortage, bruising, etc.

[0099] It can be understood that in this embodiment of the application, the defects in the target image can be used as the defects of the target object to which the target image belongs, and then the product type of the target object can be determined according to the defects of the target object. Among them, the product type here can be an OK product (i.e., a normal product) and an NG product (i.e., an NG product). An OK product means that the product is a product without defects, and an NG product means that the product is a product with defects.

[0100] Among them, the target object can be an industrial manufacturing component. In current industrial manufacturing components, the Metal Powder Injection Molding (MIM) process has very wide application scenarios, and the industrial quality requirements for MIM component finished products are relatively high. Among them, the metal powder injection molding technology is a new type of powder metallurgy near-net shaping technology formed by introducing modern plastic injection molding technology into the field of powder metallurgy. This process is widely used in many industries such as computers and their auxiliary facilities, household appliances, medical machinery parts, military parts, electrical parts, and automotive and marine parts.

[0101] Therefore, the current industrial quality inspection platform can be designed to take multi-angle photos (that is, the shooting components on multiple target terminal devices can take photos of the same component (for example, MIM components) from different angles) by photographing the frequently defective positions of the target object by the shooting components. Specifically for a certain shooting component, there will be a fixed Region of Interest (ROI) area where the photo is clear, and the rest of the area is relatively blurred and left for other shooting components to take photos. Here, the ROI area can be the clear area that the current shooting component can capture. Among them, for the same MIM component, N shooting components on N target terminal devices can take photos of the same target object to obtain N target images corresponding to the N shooting components respectively. Here, one shooting component can correspond to one target image, and N here can be a positive integer. The embodiments of the present application do not limit the specific value of N.

[0102] It can be understood that the image data processing method provided in the present application can be executed by the target terminal device, can also be executed by the server 2000, or can be jointly executed by the target terminal device and the server 2000. Among them, the embodiments of the present application are described by taking the number of target terminal devices as at least two as an example, that is, the embodiments of the present application take the above N as a positive integer greater than 1 as an example. Among them, the target terminal device may include the terminal device Z.

[0103] Among them, when the image data processing method provided in the present application is executed by the server 2000, the target terminal device (for example, the above terminal device Z) can send the target image obtained based on the shooting component to the server 2000. In this way, after the server 2000 receives the N target images provided by the target terminal device through the shooting component, it can determine the first defect output result and the second defect output result corresponding to the defect annotation areas in the N target images, and then based on the first defect output result and the second defect output result, perform multi-perspective decision analysis on the target object to which the N target images belong to obtain the object detection result corresponding to the target object.

[0104] Optionally, when the image data processing method provided in this application is executed by the target terminal device, the target terminal device may separately send the target images obtained based on the shooting component to the above-mentioned terminal device Z (terminal device Z does not need to send the target images, but still needs to obtain the target images based on the shooting component). In this way, after terminal device Z receives the (N - 1) target images provided by other terminal devices except itself, it can determine the first defect output result and the second defect output result corresponding to the defect annotation regions in the N target images, and then based on the first defect output result and the second defect output result, perform multi-perspective decision analysis on the target object to which the N target images belong to obtain the object detection result corresponding to the target object. Among them, terminal device Z can directly receive the (N - 1) target images provided by other terminal devices, or indirectly receive the (N - 1) target images forwarded by other terminal devices through server 2000.

[0105] Optionally, when the image data processing method provided in this application is jointly executed by the target terminal device and server 2000, the target terminal device (for example, terminal device Z) may separately obtain the target images based on the shooting component, determine the first defect output result and the second defect output result corresponding to the defect annotation region in the target image, and then send the first defect output result and the second defect output result determined on their respective terminal devices to server 2000. In this way, after the server receives the first defect output result and the second defect output result corresponding to the defect annotation regions in the N target images provided by the target terminal device, it can perform multi-perspective decision analysis on the target object to which the N target images belong based on the received first defect output result and the second defect output result to obtain the object detection result corresponding to the target object. Among them, if the target image includes a defect annotation region, the target terminal device may send the first defect output result and the second defect output result corresponding to the defect annotation region to server 2000. Optionally, if the target image does not include a defect annotation region, the target terminal device may not need to perform the above steps of determining the first defect output result and the second defect output result corresponding to the defect annotation region, and thus does not need to perform the step of sending the first defect output result and the second defect output result; optionally, when the target terminal device determines that the target image does not include a defect annotation region, the target terminal device may send a defect-free notice to server 2000 based on the non-inclusion of the defect annotation region to inform server 2000 that the local terminal device does not include a defect annotation region through this defect-free notice.

[0106] For ease of understanding, further, please refer to Figure 2 , Figure 2 which is a schematic diagram of a scenario for data interaction provided by an embodiment of this application. As Figure 2The server 20a shown above can be the server 2000 in the corresponding embodiment as described above. Figure 1 The terminal devices 20b and 20c shown above can be any two terminal devices in the terminal device cluster in the corresponding embodiment as described above. For ease of understanding, in the embodiments of the present application, the terminal device 3000a shown above is taken as the terminal device 20b, and the terminal device 3000c as the terminal device 20c to illustrate Figure 2 The specific process of data interaction among the server 20a, the terminal device 20b, and the terminal device 20c shown above. Figure 1 It can be understood that N terminal devices can respectively capture the same target object from different visual angles through the shooting components to obtain N target images corresponding to the N terminal devices respectively. Here, N can be a positive integer, and N equal to 2 is taken as an example for illustration. As Figure 1 shown, N (i.e., 2) terminal devices can include the terminal device 20b and the terminal device 20c. The terminal device 20b and the terminal device 20c can respectively capture the same target object from different visual angles through the shooting components to obtain target images associated with the same target object. Among them, the terminal device 20b can obtain the target image T2 of the target object, and the terminal device 20c can obtain the target image T1 of the target object. The target image T1 and the target image T2 are the N target images (i.e., 2 target images) associated with the target object. Figure 2 Furthermore, the terminal device 20b can send the captured target image T2 to the server 20a, and the terminal device 20c can send the captured target image T1 to the server 20a. In this way, after receiving the target image T1 and the target image T2, the server 20a can obtain S defect annotation regions associated with the target image T1 and the target image T2 and the first defect output results respectively corresponding to the S defect annotation regions. Here, S can be a positive integer.

[0107] As Figure 2 shown, N (i.e., 2) terminal devices can include the terminal device 20b and the terminal device 20c. The terminal device 20b and the terminal device 20c can respectively capture the same target object from different visual angles through the shooting components to obtain target images associated with the same target object. Among them, the terminal device 20b can obtain the target image T2 of the target object, and the terminal device 20c can obtain the target image T1 of the target object. The target image T1 and the target image T2 are the N target images (i.e., 2 target images) associated with the target object.

[0108] Furthermore, the terminal device 20b can send the captured target image T2 to the server 20a, and the terminal device 20c can send the captured target image T1 to the server 20a. In this way, after receiving the target image T1 and the target image T2, the server 20a can obtain S defect annotation regions associated with the target image T1 and the target image T2 and the first defect output results respectively corresponding to the S defect annotation regions. Here, S can be a positive integer.

[0109] As Figure 2As shown, the server 20a can obtain first output information 21a associated with the target image T1. The first output information 21a may include zero (i.e., 0), one, or more defect annotation regions. Here, an example is given where the first output information 21a includes 2 defect annotation regions, specifically including: defect annotation region S1 and defect annotation region S2; the server 20a can obtain first output information 21b associated with the target image T2. The first output information 21b may include zero (i.e., 0), one, or more defect annotation regions. Here, an example is given where the first output information 21b includes 1 defect annotation region, specifically including: defect annotation region S3. Among them, the defect annotation region S1, the defect annotation region S2, and the defect annotation region S3 can be collectively referred to as S defect annotation regions.

[0110] As Figure 2 shown, the first output information 21a may further include a first defect output result G1 corresponding to the defect annotation region S1 and a first defect output result G2 corresponding to the defect annotation region S2; the first output information 21b may further include a first defect output result G3 corresponding to the defect annotation region S3. Further, the server 20a can determine second defect output results corresponding to the S defect annotation regions respectively according to the S defect annotation regions (i.e., 3 defect annotation regions) and the image attribute information of the target images to which the S defect annotation regions belong.

[0111] Among them, the server 20a can obtain the image attribute information of the target image T1 to which the defect annotation region S1 belongs (for example, image attribute information X1), and determine the second defect output result G4 corresponding to the defect annotation region S1 according to the defect annotation region S1 and the image attribute information X1; the server 20a can obtain the image attribute information X1 of the target image T1 to which the defect annotation region S2 belongs, and determine the second defect output result G5 corresponding to the defect annotation region S2 according to the defect annotation region S2 and the image attribute information X1. Similarly, the server 20a can obtain the image attribute information of the target image T3 to which the defect annotation region S3 belongs (for example, image attribute information X2), and determine the second defect output result G6 corresponding to the defect annotation region S3 according to the defect annotation region S3 and the image attribute information X2.

[0112] Among them, the image attribute information X1 may include the image serial number of the target image T1, and the image serial number of the target image T1 here is determined by the camera number of the shooting component in the terminal device 20c; the image attribute information X2 may include the image serial number of the target image T2, and the image serial number of the target image T2 here is determined by the camera number of the shooting component in the terminal device 20b.

[0113] AsFigure 2 As shown, the server 20a can perform multi-perspective decision analysis on the target object based on S defect annotation regions, the first defect output results respectively corresponding to the S defect annotation regions, and the second defect output results respectively corresponding to the S defect annotation regions, so as to obtain the object detection result of the target object. In other words, the server 20a can determine the object detection result of the target object based on the defect annotation region S1, the defect annotation region S2, the defect annotation region S3, the first defect output result G1, the first defect output result G2, the first defect output result G3, the second defect output result G4, the second defect output result G5, and the second defect output result G6.

[0114] It can be seen that in the embodiment of the present application, through rough quality detection, while obtaining the defect annotation regions associated with N target images, the first defect output results corresponding to the defect annotation regions can be obtained, and then the second defect output results corresponding to the defect annotation regions can be obtained through fine quality detection. It can be understood that based on the first defect output result and the second defect output result, the quality detection of the target object corresponding to the N target images can be realized, and thus the accuracy of quality detection can be improved at the image level and the sample level.

[0115] Further, please refer to Figure 3 , Figure 3 which is a schematic flowchart of an image data processing method provided by an embodiment of the present application. This method can be executed by a server, or by a terminal device, or jointly by a server and a terminal device. The server can be the server 20a in the corresponding implementation above Figure 2 , and the terminal device can be the terminal device 20b or the terminal device 20c in the corresponding implementation above Figure 2 . For ease of understanding, the embodiment of the present application takes the execution of this method by the server as an example for description. Among them, the image data processing method may include the following steps S101 - step S103:

[0116] Step S101, obtain S defect annotation regions associated with N target images, and the first defect output results respectively corresponding to the S defect annotation regions;

[0117] Specifically, the server can obtain N target images associated with the target object and input the N target images into the instance segmentation model respectively. Among them, the N target images are obtained by N shooting components shooting the same target object respectively, and the visual angles of the N target images are different from each other. Here, N can be a positive integer. Further, the server can perform instance segmentation on the N target images through the instance segmentation model to obtain S defect annotation regions associated with the N target images and first defect output results corresponding to the S defect annotation regions respectively. Here, S can be a positive integer.

[0118] It can be understood that when the server performs instance segmentation on the N target images through the instance segmentation model, it can detect defect annotation regions in the target images or not detect defect annotation regions in the target images. In other words, the S defect annotation regions are defect annotation regions in the N target images or defect annotation regions in some of the N target images (for example, (N - 2) target images). It should be understood that each of the N target images can include zero, one or more defect annotation regions, and the embodiment of the present application does not limit the number of defect annotation regions included in each target image.

[0119] Among them, the specific process of the server performing instance segmentation on the N target images through the instance segmentation model can be referred to the description of steps S1012 - S1015 in the corresponding embodiment below. Figure 7 The description of the corresponding embodiment.

[0120] For easy understanding, please refer to Figure 4 , Figure 4 is a structural schematic diagram of a point design provided by the embodiment of the present application. As Figure 4 shown, the component 40a can be a MIM part that needs to be inspected for defects. In order to clearly capture any defects existing on the surface of the MIM part as much as possible and ensure accurate recognition by the algorithm, combined with the appearance geometric properties of the MIM part, it is necessary to set reasonable angles and lighting for photographing and imaging the MIM part.

[0121] Among them, each sample generally has many sides. For example, taking the surface detection of a certain metal device as an example, in order to cover all the appearance surfaces of the product (that is, taking into account the defects at each position and the imaging effect), the number of images taken of a sample can reach many (for example, 60+). This means that as long as one of the images is misjudged (that is, an OK defect is recognized as an NG defect, that is, overkill), then other images will also cause overkill. In addition, since there are overlapping regions in multiple images, as the number of point images increases, the probability of image overlap will continue to increase, and the overkill caused by misjudgment will show exponential growth.

[0122] As Figure 4Shown is the point design diagram of component 40a. For ease of understanding, here, it is described by taking the number of points at the point positions of component 40a as 4 as an example. The 4 point positions can specifically include: the point position corresponding to area 42a, the point position corresponding to area 42b, the point position corresponding to area 42c, and the point position corresponding to area 42d. Further, the server can generate optical images corresponding to the 4 point positions of component 40a respectively, that is, the server can generate N target images 40b associated with component 40a. Among them, the target image corresponding to area 42a can be optical image 41a, the target image corresponding to area 42b can be optical image 41b, the target image corresponding to area 42c can be optical image 41c, and the target image corresponding to area 42d can be optical image 41d.

[0123] For ease of understanding, please refer to Figure 5a and Figure 5b , Figure 5a and Figure 5b are the schematic diagrams of the architecture of a defect quality inspection scheme provided by an embodiment of the present application. As Figure 5a shown, the system architecture diagram mainly can include 3 modules. The 3 modules can specifically include: instance segmentation at the picture level implemented by the deep learning instance segmentation algorithm, fine classification at the instance level implemented by the deep learning fine classification algorithm, and multi-angle joint inference at the sample level implemented by the multi-angle joint inference algorithm. Among them, instance segmentation at the picture level and fine classification at the instance level can be collectively referred to as picture-level prediction, and multi-angle joint inference at the sample level can be collectively referred to as sample-level inference.

[0124] As Figure 5a shown, the server can perform shootings at different angles on the component (i.e., the target object) through multiple cameras (here, it is described by taking one terminal device corresponding to one camera as an example), and obtain the target images corresponding to the multiple cameras respectively. The multiple cameras can be N cameras. The N cameras can specifically include: camera O1, camera O2, …, camera O N . The target image captured by camera O1 can be target image T1 (not shown in the figure), the target image captured by camera O2 can be target image T2 (not shown in the figure), …, the target image captured by camera O N can be target image T N (not shown in the figure).

[0125] As Figure 5aAs shown, the server can input the N target images captured by the above N cameras into a deep learning instance segmentation algorithm (i.e., an instance segmentation model), and perform instance segmentation on the N target images through the deep learning instance segmentation algorithm to obtain S defect detection regions associated with the N target images, and first defect output results respectively corresponding to the S defect detection regions. Among them, the first defect output result can include the instance segmentation result respectively corresponding to each defect detection region in the S defect detection regions and the first classification information respectively corresponding to each defect detection region. It can be understood that according to the instance segmentation result respectively corresponding to each defect detection region, the defect area respectively corresponding to each defect detection region can be determined.

[0126] As Figure 5a shown, the defect types of the defect detection regions (i.e., the defect types corresponding to the deep learning instance segmentation algorithm) can include k, where k can be a positive integer, and the k defect types can specifically include: defect Y1, defect Y2,..., defect Y k . Among them, defect Y1, defect Y2,..., defect Y k can be k defect types of NG defects, and normally it can indicate that there are no defect detection regions in the target image. It should be understood that here, taking the S defect detection regions all including the above k defect types as an example for description, the S defect detection regions can include the same first classification information (i.e., k defect types).

[0127] It can be understood that the first classification information respectively corresponding to the S defect detection regions obtained by the Figure 5a deep learning instance segmentation algorithm shown above can be the above defect Y1, defect Y2,..., defect Y k , and the defect areas respectively corresponding to the S defect detection regions can be area Q1, area Q2,..., area Q k . Among them, when defect Y1, defect Y2,..., defect Y k are understood as specific defect detection regions, the defect area corresponding to defect Y1 can be area Q1, the defect area corresponding to defect Y2 can be area Q2,..., and the defect area corresponding to defect Y k can be area Q k .

[0128] As Figure 5b shown, the target image 50a can be the above Figure 5aAny one of the N target images in the corresponding embodiment (for example, target image T1). This target image 50a is the input of the above deep learning instance segmentation algorithm. Through this deep learning instance segmentation algorithm, the target image 50a after instance segmentation can be output (i.e., target image 50b). Among them, the deep learning instance segmentation algorithm can detect the defect annotation area in the target image 50a (for example, defect annotation area 200a), and display this defect annotation area 200a in the target image 50b (that is, the target image 50b is the image in the target image 50a that shows the defect annotation area 200a).

[0129] Among them, the N target images include target image L i , where i can be a positive integer less than or equal to N. The server can further execute the following steps S102 - step S103 for the target image L i . Optionally, if the server does not detect a defect annotation area in all N target images (i.e., S is equal to 0), then the server does not need to further execute the following steps S102 - step S103.

[0130] Step S102, according to the defect annotation area of the target image L i and the image attribute information of the target image L i , determine the second defect output result corresponding to the defect annotation area of the target image L i ;

[0131] Specifically, the server can determine the defect output feature corresponding to the defect annotation area of the target image L i according to the defect annotation area of the target image L i and the image serial number of the target image L i . Among them, the image attribute information of the target image L i includes the image serial number of the target image L i and the image output feature corresponding to the target image L i . Further, the server can determine the second defect output result corresponding to the defect annotation area of the target image L i according to the defect output feature corresponding to the defect annotation area of the target image L i and the image output feature corresponding to the target image L i . Among them, the server can perform fine classification processing on the defect annotation area of the target image L i through a fine classification model to determine the second defect output result corresponding to the defect annotation area of the target image L i .

[0132] Among them, the server determines the target image L iThe specific process of the defect output features corresponding to the defect annotation area can be seen in the following Figure 9 description of steps S1021 - S1022 in the corresponding embodiment.

[0133] Among them, the server determines the target image L i The specific process of the second defect output result corresponding to the defect annotation area can be seen in the following Figure 9 description of steps S1023 - S1025 in the corresponding embodiment.

[0134] It should be understood that through the instance segmentation model in the above step S101, all defects in the N target images can be ensured to be detected by the model, but it will cause serious overkill. Therefore, using the fine classification model can further refine the identification of the S defect annotation areas obtained by the instance segmentation model, that is, further subdivide the S defect annotation areas into OK defects and NG defects, thereby reducing the overkill rate. Among them, the overkill rate mainly comes from two categories: one is that some dust, foreign objects (such as hair) fall on the device, which are not defects themselves, but are very close to defects at the imaging level. For example, the instance segmentation model may misjudge hair as an NG defect; the other is that the appearance features of OK defects are very close to those of NG defects and are easily confused, so that OK defects are misjudged as NG defects. For example, dirt and bright marks themselves belong to OK defects, but they may also be detected by the instance segmentation model with a high probability and misjudged as NG defects.

[0135] Please refer to Figure 5a , the server can input the S defect detection areas obtained by the deep learning instance segmentation algorithm into the deep learning fine classification algorithm (i.e., the fine classification model), and determine the second defect output results corresponding to the S defect detection areas through the deep learning fine classification algorithm. Among them, the second defect output result may include the second classification information corresponding to each of the S defect detection areas.

[0136] As Figure 5a shown, the defect types of the defect detection areas (i.e., the defect types corresponding to the deep learning fine classification algorithm) may include (k + e), where e can be a positive integer, and the (k + e) defect types may specifically include: defect Y1, …, defect Y k , defect R1, …, defect R e . Among them, defect Y1, …, defect Y k can be k defect types of NG defects, and defect R1, …, defect R ee defect types that can be OK defects, which normally indicate that the target image does not include the defect detection area. It should be understood that here, taking the example that each of the S defect detection areas includes the above (k + e) defect types, the S defect detection areas can include the same second classification information (i.e., (k + e) defect types).

[0137] It can be understood that the second classification information corresponding to each of the S defect detection areas obtained by the Figure 5a deep learning refined classification algorithm shown can be the above-mentioned defects Y1,..., defect Y k , defect R1,..., defect R e . Among them, here, defects Y1,..., defect Y k , defect R1,..., defect R e can be understood as a specific defect detection area.

[0138] Among them, the N target images can also include the target image L j , where j can be a positive integer less than or equal to N, and the target image L j can be any target image other than the target image L i . The server can determine the second defect output result corresponding to the defect annotation area of the target image L j based on the defect annotation area of the target image L j and the image attribute information of the target image L j . Among them, for the specific process of the server determining the second defect output result corresponding to the defect annotation area of the target image L j , reference can be made to the description of determining the second defect output result corresponding to the defect annotation area of the target image L i above, and it will not be elaborated here.

[0139] Step S103, based on the first defect output results corresponding to the S defect annotation areas and the second defect output results corresponding to the S defect annotation areas, perform multi-perspective decision analysis on the target object to obtain the object detection result of the target object.

[0140] Specifically, the server can obtain the business knowledge for multi-perspective decision analysis of the target object and the target decision hyperparameters associated with the business knowledge, and generate a decision tree based on the business knowledge and the target decision hyperparameters. Further, the server can perform multi-perspective decision analysis on N target images in the decision analysis model based on the first defect output results respectively corresponding to S defect annotation regions, the second defect output results respectively corresponding to the S defect annotation regions, and the decision tree, to obtain the image detection results of the N target images respectively. Further, the server can determine the object detection result of the target object according to the image detection results of the N target images respectively.

[0141] Among them, it can be understood that the server can determine the defect detection results respectively corresponding to the S defect annotation regions in the decision analysis model based on the first defect output results respectively corresponding to the S defect annotation regions, the second defect output results respectively corresponding to the S defect annotation regions, and the decision tree. Further, the server can perform multi-perspective decision analysis on the N target images according to the defect detection results respectively corresponding to the S defect annotation regions, to obtain the image detection results of the N target images respectively.

[0142] Among them, the specific process of the server generating the decision tree can be referred to the description of steps S1031 - S1032 in the corresponding embodiment below. Figure 11 The description of steps S1031 - S1032 in the corresponding embodiment.

[0143] Among them, the specific process of the server determining the image detection results of the N target images respectively can be referred to the description of steps S1033 - S1035 in the corresponding embodiment below. Figure 11 The description of steps S1033 - S1035 in the corresponding embodiment.

[0144] Please refer to Figure 5a again. The server can input the first defect output results respectively corresponding to the S defect annotation regions and the second defect output results respectively corresponding to the S defect annotation regions into the multi-angle joint inference algorithm (i.e., the decision analysis model), and determine the defect detection results respectively corresponding to the S defect annotation regions through the multi-angle joint inference algorithm. Among them, the defect classification information in the defect detection results is determined by the first classification information respectively corresponding to the S defect annotation regions and the second classification information respectively corresponding to the S defect annotation regions. Here, the defect classification information can be defect Y1, …, defect Y k , defect R1, …, defect R e . Normal can indicate that the target image does not include a defect detection region.

[0145] Please refer to Figure 5b again. Figure 5bThe defect annotation area 200a shown is the defect annotation area 50c. The server can determine the defect detection result of the defect annotation area 50c based on the first defect output result corresponding to the defect annotation area 50c and the second defect output result corresponding to the defect annotation area 50c. Among them, the first defect output result and the second defect output result may include the defect classification corresponding to the defect annotation area 50c (i.e., the first classification information and the second classification information). The defect detection result corresponding to the defect annotation area 50c includes the defect classification information corresponding to the defect annotation area 50c, and the defect classification information can be used to indicate the defect type of the defect annotation area 50c. Among them, the defect type of the defect annotation area 50c can be Figure 5a the defects Y1 shown, …, defects Y k , defects R1, …, defects R e .

[0146] Furthermore, the server can determine the object detection result corresponding to the component according to the defect detection results corresponding to S defect annotation areas (the S defect annotation areas here can include the defect annotation area 50c) through the multi-angle joint inference algorithm. Among them, the object detection result can be used to determine whether the component is an OK product or an NG product.

[0147] For ease of understanding, please refer to Figure 6 , Figure 6 which is a schematic structural diagram of a defect quality inspection scheme provided by an embodiment of the present application. The server can input N target images associated with the target object into Figure 6 the instance segmentation model shown, and output the defect annotation area in the N target images and the first defect output result corresponding to the defect annotation area (i.e., bbox) through the instance segmentation model. Among them, the first defect output result here may include the first classification information (i.e., code), the first classification probability (i.e., score), and the instance segmentation result. Among them, according to the number of pixels indicated by the instance segmentation result, the defect area (i.e., area) can be obtained.

[0148] As Figure 6 shown, the server can input the defect annotation area into the fine classification model, and output the second defect output result corresponding to the defect annotation area through the fine classification model. Among them, the second defect output result here may include the second classification information (i.e., code2) and the second classification probability (i.e., score2). It can be understood that by combining the instance segmentation model and the fine classification model (i.e., the refined classification model), the detection and recognition at the defect instance level can be completed. However, for enterprises, what they really need in the end is the recognition at the sample level. Therefore, through the post-processing fusion strategy, the instance-level information and business knowledge can be integrated to form a sample-level judgment.

[0149] AsFigure 6 As shown, the server can input the defect annotation area, the first classification information, the second classification probability, the defect area, the second classification information, and the second classification probability into the business knowledge decision tree, and implement a multi-perspective joint decision through the business knowledge decision tree, so as to implement a sample-level inference module through this multi-perspective joint decision. Among them, the business knowledge decision tree and the multi-perspective joint decision can be collectively referred to as the post-processing fusion strategy. Through this post-processing fusion strategy, the above-mentioned target object can be determined for product-level defects, and a product-level object detection result can be obtained. Furthermore, according to the object detection result, it can be determined whether the target object is a normal product or a defective product.

[0150] It should be understood that the embodiments of the present application can detect all possible defects through an instance segmentation model with high precision (i.e., high detection rate and high recall rate), strictly prevent missed detections, and then suppress NG defects and OK defects through a fine classification model to minimize the overkill rate. Finally, the quality inspection business common sense and multi-perspective inference are combined to comprehensively decide whether the sample is an OK product or an NG product.

[0151] It should be understood that the above-mentioned instance segmentation model, fine classification model, and decision analysis model can be collectively referred to as the target network model. The instance segmentation model is obtained by iteratively training the initial instance segmentation model, and the fine classification model is obtained by iteratively training the initial instance segmentation model. Therefore, the above-mentioned initial instance segmentation model, initial fine classification model, and decision analysis model can be collectively referred to as the initial network model.

[0152] It can be seen that the embodiments of the present application can perform a rough quality inspection on N target images associated with the target object, and highly detect all defect annotation areas (i.e., S defect annotation areas) in the N target images, and then perform a fine quality inspection on the S defect annotation areas to further identify the S defect annotation areas. It can be understood that based on the first defect output result obtained from the rough quality inspection and the second defect output result obtained from the fine quality inspection, the object detection result of the decision target object can be inferred. Therefore, the embodiments of the present application can identify relatively accurate defect annotation areas and defect output results corresponding to the defect annotation areas in the N target images through rough quality inspection and fine quality inspection, and then through these accurately identified defect output results, implement the quality inspection of the target object targeted by the N target images, thereby improving the accuracy of quality inspection.

[0153] Further, please refer to Figure 7 , Figure 7 is a schematic flowchart of an image data processing method provided by the embodiments of the present application. The image data processing method may include the following steps S1011-step S1015, and steps S1011-step S1015 are Figure 3A specific embodiment of step S101 in the corresponding embodiment.

[0154] Step S1011: Obtain N target images associated with the target object, and input the N target images into the instance segmentation model respectively.

[0155] Among them, the instance segmentation model includes a feature extraction sub-network, a region prediction sub-network, and a defect recognition sub-network. It can be understood that the instance segmentation model can be used to determine S defect annotation regions associated with the N target images. The S defect annotation regions here can include M defect annotation regions in the target image L i Here, M can be a positive integer less than or equal to S, and S can be a positive integer.

[0156] It should be understood that the instance segmentation framework used by the instance segmentation model can be Mask RCNN (Mask Region Convolutional Neural Networks). The embodiments of the present application do not limit the instance segmentation framework used by the instance segmentation model.

[0157] It should be understood that the instance segmentation model is obtained after iterative training of the initial instance segmentation model. The specific process of the server iteratively training the initial instance segmentation model to obtain the instance segmentation model can be described as follows: The server can obtain defect sample annotation regions, defect sample classification information, and sample boundary regions associated with the defect sample images (i.e., images including NG defects). Further, the server can determine, in the initial instance segmentation model, the predicted defect annotation regions associated with the defect sample images, and the first predicted output result corresponding to the predicted defect annotation regions. Further, the server can determine the instance segmentation loss value of the initial instance segmentation model according to the defect sample annotation regions, defect sample classification information, sample boundary regions, predicted defect annotation regions, and the first predicted output result. Further, the server can adjust the model parameters in the initial instance segmentation model according to the instance segmentation loss value. When the adjusted initial instance segmentation model meets the model convergence condition, the adjusted initial instance segmentation model is determined as the instance segmentation model.

[0158] Among them, for the specific process of the server determining the first predicted output result corresponding to the predicted defect annotation region, reference can be made to the description of determining the first defect output result of the target image, which will not be elaborated here.

[0159] Among them, the first predicted output result may include a predicted instance segmentation result and first predicted classification information. It should be understood that the specific process by which the server determines the instance segmentation loss value of the initial instance segmentation model based on the defect sample annotation region, defect sample classification information, sample boundary region, predicted defect annotation region, and the first predicted output result can be described as follows: The server can determine the first segmentation loss value of the initial instance segmentation model based on the defect sample annotation region and the predicted defect annotation region. Further, the server can determine the second segmentation loss value of the initial instance segmentation model based on the defect sample classification information and the first predicted classification information. Further, the server can determine the third segmentation loss value of the initial instance segmentation model based on the sample boundary region and the predicted instance segmentation result. Further, the server can determine the instance segmentation loss value of the initial instance segmentation model based on the first segmentation loss value, the second segmentation loss value, and the third segmentation loss value.

[0160] It should be understood that the embodiments of the present application can consider the annotation time and cost issues of data. Both the defect sample images and the normal sample images (i.e., images including OK defects) involve defect types and defect areas. However, it is not allowed in terms of time to annotate the types and areas of all defect images. Therefore, in the embodiments of the present application, when training the initial instance segmentation model, refined annotations of the defect sample images (i.e., defect sample annotation regions, defect sample classification information, and sample boundary regions) can be used to train the initial segmentation model.

[0161] Among them, the initial instance segmentation model and the instance segmentation model can be collectively referred to as the segmentation network model. The initial instance segmentation model and the instance segmentation model are names of the segmentation network model at different times. In the training stage, the segmentation network model can be called the initial instance segmentation model, and in the prediction stage, the segmentation network model can be called the instance segmentation model.

[0162] Step S1012, input the target image L i into the feature extraction sub-network, and perform feature extraction on the target image L i through the feature extraction sub-network to obtain multi-resolution features corresponding to the target image L i ;

[0163] Specifically, the server can input the target image L i into the feature extraction sub-network, and perform feature extraction on the target image L i through the feature extraction sub-network to obtain auxiliary image features corresponding to at least two resolutions. Among them, one resolution corresponds to one or more auxiliary image features. Further, the server can perform feature aggregation on the auxiliary image features with the same resolution to obtain aggregated auxiliary image features. Further, the server can perform feature extraction on the aggregated auxiliary image features to obtain the target image Li The corresponding candidate image features. Further, the server may generate a target image L based on the auxiliary image features and candidate image features corresponding to at least two resolutions respectively i The corresponding multi-resolution features.

[0164] Among them, the feature extraction sub-network may include a feature extraction network layer, and the server may perform feature extraction on the target image L through the feature extraction network layer i Or aggregate the auxiliary image features for feature extraction. Here, the feature extraction may be upsampling processing or downsampling processing. It should be understood that the feature extraction network layer here may be a CNN (Convolutional Neural Network), and the convolutional neural network may perform a convolution operation on the target image L i Or aggregate the auxiliary image features. The embodiments of the present application do not limit the type of the feature extraction network layer.

[0165] Among them, the way for the server to perform feature aggregation on the auxiliary image features with the same resolution may be the way of feature splicing, or the way of feature addition, or the way of feature weighted average. The embodiments of the present application do not limit the specific way of feature fusion.

[0166] Among them, the feature extraction sub-network may further include a feature fusion network layer, and the server may fuse the auxiliary image features and candidate image features corresponding to at least two resolutions respectively into the multi-resolution features corresponding to the target image L through the feature fusion network layer i It should be understood that the feature fusion network layer here may be an FPN network (Feature Pyramid Networks). The embodiments of the present application do not limit the type of the feature fusion network layer.

[0167] It should be understood that the feature extraction sub-network may be an HRNet network (High Resolution Net). The embodiments of the present application do not limit the type of the feature extraction sub-network. Among them, the HRNet network can maintain high-resolution features and can fully fuse multi-resolution features, thereby improving the small defect detection performance. Among them, the feature extraction sub-network may be obtained by any combination of one or more feature extraction network layers. The embodiments of the present application do not limit the structure of the feature extraction network layer in the feature extraction sub-network.

[0168] Step S1013, input the multi-resolution features corresponding to the target image L i Into the region prediction sub-network, and perform region prediction on the multi-resolution features corresponding to the target image L through the region prediction sub-network to obtain the target image L i The corresponding multi-resolution features, and obtain the target image Li M object regions to be predicted therein;

[0169] Among them, the region prediction sub-network can directly predict the object regions to be predicted and determine the positions of the object regions to be predicted. It should be understood that the region prediction sub-network can be an RPN network (Region Proposal Network, candidate region network), and the embodiments of the present application do not limit the type of the region prediction sub-network.

[0170] Step S1014, input the M object regions to be predicted and the multi-resolution features corresponding to the target image L i into the defect recognition sub-network, and perform defect recognition on the M object regions to be predicted and the multi-resolution features corresponding to the target image L through the defect recognition sub-network i to obtain the instance segmentation results corresponding to the M defect annotation regions respectively, the first classification probabilities corresponding to the M defect annotation regions respectively, and the first classification information corresponding to the M defect annotation regions respectively;

[0171] Specifically, the server can input the M object regions to be predicted and the multi-resolution features corresponding to the target image L i into the defect recognition sub-network, and map the M object regions to be predicted to the multi-resolution features corresponding to the target image L through the defect recognition sub-network i to obtain the candidate region features corresponding to the M object regions to be predicted respectively. Further, the server can perform feature alignment on the M candidate region features to obtain the aligned region features corresponding to the M candidate region features respectively. Further, the server can perform a convolution operation on the M aligned region features to obtain the classification region features corresponding to the M aligned region features respectively and the segmentation region features corresponding to the M aligned region features respectively. Further, the server can perform a full connection operation on the M classification region features to determine the region features corresponding to the M aligned region features respectively and the classification features corresponding to the M aligned region features respectively. Based on the M region features, determine the M defect annotation regions (i.e., detection frames). Based on the M classification features, determine the first classification probabilities corresponding to the M defect annotation regions respectively and the first classification information corresponding to the M defect annotation regions respectively. Further, the server can perform a convolution operation on the M segmentation region features to determine the segmentation features corresponding to the M aligned region features respectively. Based on the M segmentation features, determine the instance segmentation results corresponding to the M defect annotation regions respectively (i.e., pixel-level prediction).

[0172] Among them, the candidate region feature is a feature associated with the object region to be predicted intercepted on the multi-resolution feature, that is, the server can determine the candidate region feature in the multi-resolution feature according to the position of the object region to be predicted.

[0173] Among them, the defect recognition sub-network may include a feature alignment network layer, and the server may align the candidate region features to the same feature dimension through the feature alignment network layer. It should be understood that the feature alignment network layer here may be ROIAlign, and the feature alignment network layer here may also be ROIPooling. The embodiments of the present application do not limit the type of the feature alignment network layer.

[0174] Among them, the defect recognition sub-network may further include a fully convolutional network (Fully Convolutional Network, abbreviated as FCN). The server may perform a convolution operation on the aligned region features and the segmentation region features through the fully convolutional network. Among them, the defect recognition sub-network may further include a classification fully connected layer and a region fully connected layer. By performing a fully connected operation on the classification region features through the classification fully connected layer, classification features may be obtained. By performing a fully connected operation on the classification region features through the region fully connected layer, region features may be obtained.

[0175] Step S1015: Use the instance segmentation results corresponding to the M defect annotation regions, the first classification probabilities corresponding to the M defect annotation regions, and the first classification information corresponding to the M defect annotation regions as the first defect output results corresponding to the M defect annotation regions.

[0176] Among them, the first classification probability may represent the probability that the defect annotation region belongs to the first classification information. Each defect annotation region has a classification probability corresponding to all classification information. The first classification probability is the maximum classification probability among these classification probabilities, and the first classification information is the classification information corresponding to the first classification probability.

[0177] For ease of understanding, please refer to Figure 8 , Figure 8 which is a schematic structural diagram of an instance segmentation model provided by an embodiment of the present application. It can be understood that when the structural diagram shown in Figure 8 corresponds to the structural diagram of the instance segmentation model, Figure 8 the image 80a shown in

[0178] As Figure 8 shown, the server may input the image 80a into the feature extraction sub-network 80b, output the multi-resolution features corresponding to the image 80a through the feature extraction sub-network 80b, and then input the multi-resolution features corresponding to the image 80a into the region prediction sub-network, and output M to-be-predicted object regions in the image 80a through the region prediction sub-network. Among them, M here may be a positive integer.

[0179] Among them, it can be understood that by extracting features from the image 80a through the feature extraction sub-network 80b, the feature 81a can be obtained, and then by extracting features from the feature 81a, the feature 82a can be obtained. Further, the server can extract features from the feature 81a to obtain a first auxiliary image feature (not shown in the figure), and extract features from the feature 82a to obtain a second auxiliary image feature (not shown in the figure). Among them, the first auxiliary image feature and the second auxiliary image feature have the same resolution. Further, the server can perform feature fusion on the first auxiliary image feature and the second auxiliary image feature to obtain the feature 83a.

[0180] Further, the server can use the feature 81a as the feature 81b, the feature 82a as the feature 82b, and the feature 83a as the feature 83b. Further, the server can use the feature 81b as the feature 81c, the feature 82b as the feature 82c, and the feature 83b as the feature 83c. Among them, the feature 81c and the feature 82c can be called auxiliary image features, and the feature 83c can be called an aggregated auxiliary image feature. Further, the server can generate multi-resolution features corresponding to the image 80a based on the feature 81c, the feature 82c, and the feature 83c. Optionally, the server can also perform further feature extraction on the feature 83c to obtain a candidate image feature (not shown in the figure), and then generate multi-resolution features corresponding to the image 80a based on the feature 81c, the feature 82c, and the candidate image feature.

[0181] As Figure 8 shown, the server can input the multi-resolution features corresponding to the image 80a and the M regions of objects to be predicted into the defect recognition sub-network 80c, and determine, through the defect recognition sub-network 80c, the instance segmentation results respectively corresponding to the M defect annotation regions, the first classification probabilities respectively corresponding to the M defect annotation regions, and the first classification information respectively corresponding to the M defect annotation regions. Among them, the defect recognition sub-network 80c can include a feature alignment network layer 84a, a fully convolutional network 84b, and a fully convolutional network 84c.

[0182] Among them, it can be understood that the server can input the multi-resolution features and the M regions of objects to be predicted into the feature alignment network layer 84a, and output, through the feature alignment network layer 84a, the aligned region features corresponding to the image 80a, and then input the aligned region features into the fully convolutional network 84b, and output, through the fully convolutional network 84b, the classification region features and the segmentation region features corresponding to the aligned region features. As Figure 8As shown, the server can input the segmented region features into the fully convolutional network 84c, and determine the instance segmentation result corresponding to the image 80a through the fully convolutional network 84c; the server can determine the region features (not shown in the figure) and the classification features 84d based on the classified region features, and then determine the defect annotation region corresponding to the region features through the defect recognition sub-network 80c, and determine the first classification information and the classification probability corresponding to the classification features 84d through the defect recognition sub-network 80c. Among them, the defect annotation region determined by the defect recognition sub-network 80c is Figure 8 the region 80d shown, and this region 80d is an instance corresponding to the instance segmentation model.

[0183] It should be understood that the embodiments of the present application can realize data perturbation, brightness and contrast, rotation and displacement, multi-scale, etc. in data augmentation to expand the diversity of online data and ensure the stability of the model in the production environment. In other words, the embodiments of the present application can perform data augmentation on the sample images (for example, defect sample images, normal sample images) during the model training process, and use the data-augmented sample images to train the model to improve the model's ability to adapt to different types of data. For example, adjust the brightness of the sample images, rotate the sample images, and scale the sample images. It should be understood that when performing enhancement operations such as rotation and scaling on the sample images, the annotation boxes in the sample images will change accordingly.

[0184] It should be understood that the embodiments of the present application can also construct a three-level label system that combines physical labels and image labels in the label system (including the defect type under the microscope, whether the defect is judged as a defect or non-defect on the production line, and whether the defect is clearly distinguishable at the image level), which can ensure the accuracy of the training data annotation to a certain extent. Among them, due to reasons such as light and shooting angle, the imaging at the image level will be unclear. Among them, the defect type is a fine-grained index of the defect, that is, the classification information of the defect. For example, bruise; the production line judgment is a coarse-grained index of the defect, and the defect can be divided into OK defects and NG defects; the image level represents the index of the defect at the display level. It can be understood that under the influence of light, angle, etc., the defect is prone to be unclear (for example, there is originally a defect, but under a certain lighting condition, the defect is not photographed); under the condition that the defect is clear, due to various factors, the imaging of the defect will be unclear; under the condition that the defect is clear, the imaging is clearly distinguishable, and then the type of the defect can be determined.

[0185] It should be understood that the label system at the image level has multiple uses. In the initial stage of the device, samples are collected and imaged by the device, and then through comparison with the physical object for evaluation, an index for the clarity of physical defects can be determined. If it is determined that the index for physical clarity meets the index conditions (for example, the index is greater than 95%), it is determined that there is no problem with the device imaging; otherwise, the device needs to be further debugged. In the case of clear imaging, if there are still unclear samples, the clear samples can be labeled and the unclear samples can be deleted to improve the training ability of the samples for the model. After the model is trained, to evaluate whether the model meets the expectations, the index of the model can be determined through the false negative rate, and here the false negative rate will be affected by the clarity of the image.

[0186] Optionally, when Figure 8 the shown structural schematic diagram corresponds to the structural schematic diagram of the initial instance segmentation model, Figure 8 the shown image 80a can be a defective sample image. Through Figure 8 the shown structural schematic diagram, the predicted defective annotation region, the predicted instance segmentation result, and the first predicted classification information associated with the defective sample image can be determined. Among them, the specific process by which the server determines the predicted defective annotation region, the predicted instance segmentation result, and the first predicted classification information associated with the defective sample image through Figure 8 the shown structural schematic diagram can refer to the description of determining the defective annotation region, the instance segmentation result, and the first classification information associated with the target image (i.e., image 80a) through Figure 8 the shown structural schematic diagram above, and will not be elaborated here.

[0187] It can be seen that the instance segmentation model in the embodiments of the present application can perform instance segmentation on N target images, obtain S defective detection regions associated with the N target images, and the first defective output results corresponding to the S defective detection regions respectively, thereby realizing a rough quality detection of the accuracy of the N target images. It can be understood that through the instance segmentation model, all defects can be detected as much as possible in the N target images, the defects in the N target images are labeled at the instance level, the defect type (i.e., the first classification information), the defect rectangle box, and the external polygon of the defect are output, and then the pixel area corresponding to the defect is predicted, so as to fully ensure a low false negative rate.

[0188] Furthermore, please refer to Figure 9 , Figure 9 which is a flowchart of an image data processing method provided by the embodiments of the present application. The image data processing method may include the following steps S1021 - step S1025, and steps S1021 - step S1025 are Figure 3 a specific embodiment of step S102 in the corresponding embodiment.

[0189] Step S1021: Determine the target image L i to obtain the regional coordinates of the defect annotation area of i . Based on the regional coordinates and the image serial number of the target image L i generate the defect input features corresponding to the defect annotation area of the target image L i and input the defect input features into the fine classification model;

[0190] Among them, the fine classification model includes a perception sub-network and a feature recognition sub-network.

[0191] It can be understood that the judgment of the defect type is closely related to the point position and the position where the defect occurs. Therefore, in the model establishment stage, the server can concatenate the coordinate information of the defect position, that is, the coordinates of the defect bounding box (i.e., the defect detection box) (the abscissa and ordinate of the upper left position of the bounding box, and the abscissa and ordinate of the lower right position of the bounding box) and the point position ordinal number coding into a vector, and use the vector obtained by this coding as the defect input feature.

[0192] It should be understood that the fine classification model is obtained after iterative training of the initial fine classification model. The specific process of the server iteratively training the initial fine classification model to obtain the fine classification model can be described as follows: The server can obtain the defect sample annotation area and defect sample classification information associated with the defect sample image, and obtain the normal sample annotation area and normal sample classification information associated with the normal sample image. Further, the server can, in the initial fine classification model, determine the second predicted output result corresponding to the defect sample annotation area according to the image attribute information of the defect sample annotation area and the defect sample image, and determine the first classification loss value of the initial fine classification model according to the second predicted output result corresponding to the defect sample annotation area and the defect sample classification information. Further, the server can determine the second predicted output result corresponding to the normal sample annotation area according to the image attribute information of the normal sample annotation area and the normal sample image, and determine the second classification loss value of the initial fine classification model according to the second predicted output result corresponding to the normal sample annotation area and the normal sample classification information. Further, the server can determine the fine classification loss value of the initial fine classification model according to the first classification loss value and the second classification loss value. Further, the server can adjust the model parameters in the initial fine classification model according to the fine classification loss value, and when the adjusted initial fine classification model meets the model convergence condition, determine the adjusted initial fine classification model as the fine classification model.

[0193] Among them, for the specific process of the server obtaining the image attribute information of the defect sample image and the image attribute information of the normal sample image, reference can be made to the description of obtaining the image attribute information of the target image in the corresponding embodiment above Figure 3 and details will not be elaborated here.

[0194] Among them, for the specific process of the server to determine the second prediction output result corresponding to the normal sample annotation area and the specific process of determining the second prediction output result corresponding to the normal sample annotation area, reference can be made to the description of determining the second defect output result of the target image, which will not be elaborated here.

[0195] Among them, the second prediction output result corresponding to the defective sample annotation area may include the second prediction classification information of the defective sample annotation area. It should be understood that the specific process for the server to determine the first classification loss value of the initial fine classification model based on the second prediction output result corresponding to the defective sample annotation area and the defective sample classification information can be described as: the server can determine the first classification loss value of the initial fine classification model according to the second prediction classification information of the defective sample annotation area and the defective sample classification information.

[0196] Among them, the second prediction output result corresponding to the normal sample annotation area may include the second prediction classification information of the normal sample annotation area. It should be understood that the specific process for the server to determine the second classification loss value of the initial fine classification model based on the second prediction output result corresponding to the normal sample annotation area and the normal sample classification information can be described as: the server can determine the second classification loss value of the initial fine classification model according to the second prediction classification information of the normal sample annotation area and the normal sample classification information.

[0197] It should be understood that the embodiments of the present application can greatly alleviate the problem of slow annotation of detection frames (i.e., bboxes) and boundary polygons (i.e., masks). First, this framework (i.e., the instance segmentation model and the fine classification model) only needs to finely annotate NG defects (defect types, detection frames, boundary polygons of defects), and roughly annotate OK defects (defect types, detection frames); second, by training the instance segmentation model with the finely annotated NG defects, the detection rate and the accuracy of the predicted area of defects of the instance segmentation model can meet the requirements of the production line. Furthermore, the finely marked NG defects and the roughly marked OK defects can be used to train the fine classification model. Based on this, the training data annotation method of the instance segmentation model and the fine classification model provided by the embodiments of the present application significantly reduces the cost of data annotation.

[0198] Among them, the detection rate is the recall rate, and the recall rate indicates how many NG products in the sample are predicted correctly; the accuracy rate indicates how many products' areas of NG products and OK products in the sample are predicted correctly. It can be understood that the server can determine the intersection over union (IoU) between the predicted defect area and the detection frame. When the IoU is greater than the area threshold, it is determined that the predicted defect area of the defect annotation area is predicted accurately. The specific value of the area threshold is not limited here.

[0199] Among them, the initial fine-grained classification model and the fine-grained classification model can be collectively referred to as the classification network model. The initial fine-grained classification model and the fine-grained classification model are the names of the classification network model at different times. In the training stage, the classification network model can be called the initial fine-grained classification model, and in the prediction stage, the classification network model can be called the fine-grained classification model.

[0200] Step S1022: Perform a fully connected operation on the defect input features through the perception sub-network to determine the target image L i The defect output features corresponding to the defect annotation area of;

[0201] Among them, the perception sub-network may include one or more fully connected layers connected in series, and multiple fully connected layers can achieve non-linear classification. It should be understood that the embodiments of the present application do not limit the number of fully connected layers in the perception sub-network.

[0202] Step S1023: Input the target image L i Into the feature recognition sub-network, and perform feature recognition on the target image L through the feature recognition sub-network i To obtain the image output features corresponding to the target image L i ;

[0203] It should be understood that the feature recognition sub-network can be a ResNet network (Deep Residual Network) of the Attention mechanism (i.e., the attention mechanism). The embodiments of the present application do not limit the type of the feature recognition sub-network. Among them, adopting the ResNet architecture and adding the attention mechanism to the structure of the refined classification model (i.e., the fine-grained classification model) can further perform fine-grained discrimination from the defect detail information.

[0204] Among them, the attention mechanism used in the embodiments of the present application can be a CBAM module (Convolutional Block Attention Module). This CBAM module is an attention mechanism based on convolutional blocks, which can fuse spatial attention and channel attention. The embodiments of the present application do not limit the type of the attention mechanism.

[0205] Step S1024: Perform feature fusion on the defect output features corresponding to the defect annotation area of the target image L i And the image output features corresponding to the target image L i To obtain the fusion output features corresponding to the defect annotation area of the target image L i ;

[0206] It should be understood that the server pairs the defect output features corresponding to the defect annotation area of the target image L i And the target image Li The way of fusing the corresponding image output features can be the way of feature splicing, the way of feature addition, or the way of feature weighted average. The embodiments of the present application do not limit the specific way of feature fusion.

[0207] Step S1025, based on the target image L i and the classifier of the fine classification model, determine the second defect output result corresponding to the defect annotation area of the target image L i

[0208] Specifically, the server may input the fusion output feature corresponding to the defect annotation area of the target image L i into the classifier of the fine classification model, and determine the matching degree between the fusion output feature corresponding to the defect annotation area of the target image L i and the sample output feature in the classifier. Among them, the matching degree is used to describe the probability that the defect annotation area of the target image L i belongs to the sample classification label corresponding to the sample output feature. Here, the classifier may be a fully connected layer (the fully connected layer is a non-linear classifier). Further, the server may use the sample classification label corresponding to the sample output feature with the maximum matching degree as the second classification information corresponding to the defect annotation area of the target image L i and use the maximum matching degree as the second classification probability corresponding to the defect annotation area of the target image L i Further, the server may use the second classification information corresponding to the defect annotation area of the target image L i and the second classification probability corresponding to the defect annotation area of the target image L i as the second defect output result corresponding to the defect annotation area of the target image L i

[0209] Among them, the second classification probability may represent the probability that the defect annotation area belongs to the second classification information. Each defect annotation area has a classification probability corresponding to all classification information. The second classification probability is the maximum classification probability among these classification probabilities, and the second classification information is the classification information corresponding to the second classification probability.

[0210] For ease of understanding, please refer to Figure 10 Figure 10 which is a schematic structural diagram of a fine classification model provided by the embodiments of the present application. It can be understood that when the structural diagram shown in Figure 10 corresponds to the structural diagram of the fine classification model, the fine classification model 100c may include a feature recognition sub-network and a perception sub-network. Among them, Figure 10 the image 100a shown may be any one of the N target images, such as​​​Figure 10 The defect annotation area 100b shown can be any defect annotation area in the image 100a.

[0211] As Figure 10 shown, the server can input the image 100a into the feature recognition sub-network, and output the image output features corresponding to the image 100a through the feature recognition sub-network; the server can input the defect annotation area 100b and the image serial number of the image 100a into the perception sub-network, and output the defect output features corresponding to the defect annotation area 100b through the perception sub-network. Further, the server can perform feature fusion on the image output features and the defect output features to obtain the fusion output feature 100d corresponding to the defect annotation area 100b, and then input the fusion output feature 100d into the classifier in the fine classification model 100c, and output the second defect output result corresponding to the defect annotation area 100b through the classifier.

[0212] Among them, the defect output feature can be a multi-dimensional vector. For example, the multi-dimension here can be 256 dimensions. Among them, the image output feature can be a multi-dimensional vector. For example, the multi-dimension here can be 512 dimensions.

[0213] Optionally, when the structure diagram shown Figure 10 corresponds to the structure diagram of the initial fine classification model, Figure 10 the image 100a shown can be a defect sample image or a normal sample image. Through Figure 10 the structure diagram shown, the second prediction classification information associated with the defect sample image and the second prediction classification information associated with the normal sample image can be determined. Among them, the specific process by which the server determines the second prediction classification information associated with the defect sample image and the second prediction classification information associated with the normal sample image through Figure 10 the structure diagram shown can refer to the description of determining the second classification information associated with the target image (i.e., the image 100a) through Figure 10 the structure diagram shown above, which will not be elaborated here.

[0214] It can be seen that the fine classification model in the embodiments of the present application can determine the image output features corresponding to the target image, as well as the defect output features corresponding to the defect annotation area of the target image, and then determine the second defect output result corresponding to the defect annotation area of the target image according to the image output features corresponding to the target image and the defect output features corresponding to the defect annotation area of the target image, thereby realizing the fine quality detection of S defect annotation areas. It can be understood that through the fine classification model, the S defect annotation areas can be refined and divided, and the defect types (i.e., the second classification information) of the S defects can be output, and then the pseudo-defects in the S defect annotation areas can be identified, thereby reducing the overkill rate.

[0215] Further, please refer to Figure 11 , Figure 11 which is a schematic flowchart of an image data processing method provided by an embodiment of the present application. The image data processing method may include the following steps S1031 - step S1036, and steps S1031 - step S1036 are Figure 3 a specific embodiment of step S103 in the corresponding embodiment.

[0216] Step S1031: Obtain business knowledge and a hyperparameter search model for multi - perspective decision - making analysis of a target object, and generate a set of hyperparameters associated with the business knowledge through the hyperparameter search model;

[0217] Among them, the set of hyperparameters includes one or more groups of decision hyperparameters; each group of decision hyperparameters in one or more groups of decision hyperparameters includes one or more hyperparameters; one or more groups of decision hyperparameters are used to balance at least two evaluation indicators corresponding to the decision - making analysis model. For example, the at least two evaluation indicators may be the over - kill rate and the missed - detection rate. Among them, the over - kill rate and the missed - detection rate are for the sample level.

[0218] It can be understood that a confusion matrix can be used to represent the number of predicted categories (i.e., columns) and actual categories (i.e., rows). The confusion matrix output after passing through the quality inspection system can be seen in Table 1 below:

[0219] Table 1

[0220]

[0221] Among them, TP (True Positive) can represent the number of OK products predicted as OK products, FN (False Negative) represents the number of OK products predicted as NG products, FP (False Positive) represents the number of NG products predicted as OK products, and TN (True Negative) represents the number of NG products predicted as NG products. It can be understood that through TP, FN, FP, and TN, the over - kill rate and the missed - detection rate of the product can be determined. Here, not all evaluation indicators are listed one by one.

[0222] Among them, the over - kill rate is the ratio of the system misjudging (i.e., misclassifying) OK products as NG products. The over - kill rate can be seen in the following formula (1):

[0223]

[0224] Among them, the missed - detection rate is the ratio of the system misjudging NG products as OK products. The missed - detection rate can be seen in the following formula (2):

[0225]

[0226] It can be understood that the overkill rate and the missed detection rate can be determined as a pair of conflicting indicators through the above formulas (1) and (2). For the same model, strictly controlling the missed detection rate will necessarily cause the overkill rate to increase; strictly controlling the overkill rate will necessarily cause the missed detection rate to increase.

[0227] Among them, in the quality inspection industry of MIM parts, the judgment of OK products and NG products depends on a lot of business knowledge and comprehensive judgment of multiple positions. The main business knowledge is listed as follows: (1) Cracks have the strictest assessment requirements, and the missed detection index is close to zero. The defects and non-defects of crack judgment have nothing to do with the area; (2) Defects have sizes and depths. For example, some small and shallow defects can be regarded as OK (such as pressing injuries, etc.). Therefore, the judgment of defects and non-defects of some defects needs to consider the area. Defects smaller than the given area can be regarded as OK defects, and those larger than the given area can be regarded as NG defects; (3) Some defects, although they look like NG defects in appearance, belong to OK defects, such as bright imprints and dirt; (4) Each position picture of the sample has its own ROI (i.e., region of interest), and there is a large overlap between the pictures of different positions. If each position is only responsible for the detection result of its own ROI, the overkill rate can be effectively reduced; (5) Among the multiple positions corresponding to the sample, as long as 1 position is judged as an NG defect, then this sample belongs to an NG sample; if all positions are judged as OK defects, then the sample belongs to an OK sample.

[0228] For ease of understanding, please refer to Figure 12 , Figure 12 which is a schematic flowchart of a process for generating a set of hyperparameters provided by an embodiment of the present application. As Figure 12 is a schematic flowchart of the process for the hyperparameter search model to perform decision hyperparameter search. The hyperparameter search model can be the Pareto optimal algorithm, then Figure 12 shows the hyperparameter optimization process of the Pareto optimal algorithm. Among them, the optimization variables of the Pareto optimal algorithm are decision tree hyperparameters based on business knowledge (i.e., decision hyperparameters).

[0229] As Figure 12 shown, the server can execute step S21 and step S22 to generate an initial population P. Here, the initial population P can be the optimal solution at the current moment. Further, the server can execute step S23 to perform operations (such as crossover, mutation, and selection) on the initial population P through an evolutionary algorithm (EA) to obtain a new population R. Among them, more optimal solutions can be obtained through continuous evolutionary operations. Further, the server can execute step S24, and through step S24, a non-dominated set (i.e., Nset) of the initial population P and the new population R (i.e., PUR) can be constructed.

[0230] Among them, when designing the Pareto optimal algorithm, a size threshold of the non-dominated set Nset is set. If the size of the current non-dominated set Nset is greater than or equal to the size threshold, then in step S25, the non-dominated set Nset needs to be adjusted according to a certain strategy (that is, the scale of the non-dominated set Nset is adjusted). It can be understood that the adjustment in step S25 can, on the one hand, make the non-dominated set Nset meet the size requirement, and on the other hand, make the non-dominated set Nset meet the distribution requirement.

[0231] Furthermore, the server can execute step S26 to determine whether the non-dominated set Nset meets the termination condition. If the non-dominated set Nset meets the termination condition, then step S28 is executed, and in step S28, the non-dominated set result (that is, the Pareto optimal solution, the optimal solution set) is output. Optionally, if the non-dominated set Nset does not meet the termination condition, then step S27 is executed (at this time, P is less than or equal to the non-dominated set), and then step S23 is executed to use the non-dominated set as the new initial population P and perform an evolutionary operation on the new initial population P, that is, copy the individuals in the non-dominated set Nset into P and continue the next round of evolution. Among them, the termination condition can be an iteration count limit or an iteration transformation limit. The iteration count limit means iterating a specified number of times, and the iteration transformation limit indicates that the non-dominated set has not changed after multiple iterations. The embodiments of the present application do not limit the termination condition here.

[0232] Among them, the above evolutionary algorithm can be a heuristic algorithm or an evolutionary algorithm. Using a heuristic algorithm or an evolutionary algorithm to solve this problem is a good idea. It can be understood that the advantage of the heuristic algorithm is that it does not need to know the specific form of the objective function, has no requirements for the differentiability and derivability of the objective function, and supports multi-objective optimization; with the in-depth research, the theory of the evolutionary algorithm has gradually tended to be mature and perfect. Among them, many evolutionary algorithms represented by the genetic algorithm have the characteristics of generating multiple points and performing multi-directional searches, so they are very suitable for solving multi-objective optimization problems with a very complex search space for this kind of optimal solution.

[0233] Step S1032, obtain the target decision hyperparameters that meet the hyperparameter acquisition conditions from the hyperparameter set, and generate a decision tree according to the business knowledge and the target decision hyperparameters;

[0234] It can be understood that the above business knowledge can be simply summarized as decision trees. By comprehensively using these decision trees, it is possible to infer OK / NG at the sample level. However, these decision trees include many hyperparameters. For example, for a single instance, the detection threshold and area of the instance segmentation model, the detection threshold of the classification model, etc. For the same model, different combinations of hyperparameters may correspond to different missed detection rates and overkill rates. Selecting hyperparameters that meet the business indicators can have an important impact on the comprehensive performance of the model. Therefore, the embodiments of the present application can select appropriate target decision hyperparameters for the missed detection rate and overkill rate (i.e., hyperparameter acquisition conditions).

[0235] Among them, the target decision hyperparameters include instance segmentation hyperparameters, segmentation area hyperparameters, and fine classification hyperparameters. The instance segmentation hyperparameter is the above-mentioned detection threshold of the segmentation model, the segmentation area hyperparameter is the above-mentioned detection area of the instance segmentation model, and the fine classification hyperparameter is the above-mentioned detection threshold of the classification model.

[0236] Step S1033: Obtain the instance segmentation results corresponding to the S defect annotation regions from the first defect output results corresponding to the S defect annotation regions respectively, and determine the defect region areas corresponding to the S defect annotation regions respectively according to the instance segmentation results corresponding to the S defect annotation regions respectively.

[0237] Step S1034: Obtain the first classification probabilities corresponding to the S defect annotation regions respectively, and the first classification information corresponding to the S defect annotation regions respectively, from the first defect output results corresponding to the S defect annotation regions respectively, and obtain the second classification probabilities corresponding to the S defect annotation regions respectively, and the second classification information corresponding to the S defect annotation regions respectively, from the second defect output results corresponding to the S defect annotation regions respectively.

[0238] Step S1035: In the decision analysis model, perform multi-perspective decision analysis on the N target images according to the first classification information corresponding to the S defect annotation regions respectively, the second classification information corresponding to the S defect annotation regions respectively, the first classification probabilities corresponding to the S defect annotation regions respectively, the second classification probabilities corresponding to the S defect annotation regions respectively, the defect region areas corresponding to the S defect annotation regions respectively, the S defect annotation regions, and the instance segmentation hyperparameters, segmentation area hyperparameters, and fine classification hyperparameters indicated by the decision tree, to obtain the image detection results of the N target images respectively.

[0239] Specifically, the server can, in the decision analysis model, determine the defect detection results corresponding to the S defect annotation regions based on the first classification information corresponding to the S defect annotation regions respectively, the second classification information corresponding to the S defect annotation regions respectively, the first classification probabilities corresponding to the S defect annotation regions respectively, the second classification probabilities corresponding to the S defect annotation regions respectively, the defect region areas corresponding to the S defect annotation regions respectively, the S defect annotation regions, as well as the instance segmentation hyperparameters, segmentation area hyperparameters, and fine classification hyperparameters indicated by the decision tree. Further, the server can perform multi-perspective decision analysis on the N target images based on the defect detection results corresponding to the S defect annotation regions respectively, and obtain the image detection results of the N target images respectively.

[0240] Among them, the server can, in the decision analysis model, determine the first detection results corresponding to the S defect annotation regions respectively based on the first classification information corresponding to the S defect annotation regions respectively, the first classification probabilities corresponding to the S defect annotation regions respectively, and the instance segmentation hyperparameters indicated by the decision tree. Further, the server can determine the second detection results corresponding to the S defect annotation regions respectively based on the second classification information corresponding to the S defect annotation regions respectively, the second classification probabilities corresponding to the S defect annotation regions respectively, and the fine classification hyperparameters indicated by the decision tree. Further, the server can determine the third detection results corresponding to the S defect annotation regions respectively based on the defect region areas corresponding to the S defect annotation regions respectively and the segmentation area hyperparameters indicated by the decision tree. Further, the server can determine the defect detection results corresponding to the S defect annotation regions respectively based on the first detection results corresponding to the S defect annotation regions respectively, the second detection results corresponding to the S defect annotation regions respectively, and the third detection results corresponding to the S defect annotation regions respectively. Further, the server can perform multi-perspective decision analysis on the N target images based on the defect detection results corresponding to the S defect annotation regions respectively, and obtain the image detection results of the N target images respectively.

[0241] Among them, if the first classification probability is greater than the instance segmentation hyperparameter, it indicates that the defect annotation region is the defect type corresponding to the first classification information; if the first classification probability is less than or equal to the instance segmentation hyperparameter, it indicates that the defect annotation region is not the defect type corresponding to the first classification information.

[0242] Among them, if the second classification probability is greater than the fine classification hyperparameter, it indicates that the defect annotation region is the defect type corresponding to the second classification information; if the second classification probability is less than or equal to the fine classification hyperparameter, it indicates that the defect annotation region is not the defect type corresponding to the second classification information.

[0243] Among them, if the area of the defect region is greater than the segmentation area hyperparameter, it indicates that the defect annotation region can be an NG defect. If the area of the defect region is less than or equal to the segmentation area hyperparameter, it indicates that the defect annotation region can be an OK defect.

[0244] Among them, the N target images may include the target image L i , and here the target image L i is taken as an example for illustration. The server can determine the image detection result of the target image L i according to the defect detection results of the M defect annotation regions in the target image L i . Among them, the defect detection result can indicate whether the defect annotation region is a defect or not. Among them, if there is no NG defect (that is, there may be an OK defect, or there is neither an OK defect nor an NG defect) in the defect annotation region of the target image L i , it is determined that the target image L i is a non-defective image; if there is an NG defect (there may also be an OK defect, or there is no OK defect) in the defect annotation region of the target image L i , it is determined that the target image L i is a defective image.

[0245] Optionally, the server can determine the first detection results of the N target images respectively in the decision analysis model according to the first classification information corresponding to the S defect annotation regions respectively, the first classification probabilities corresponding to the S defect annotation regions respectively, and the instance segmentation hyperparameters indicated by the decision tree. Further, the server can determine the second detection results of the N target images respectively according to the second classification information corresponding to the S defect annotation regions respectively, the second classification probabilities corresponding to the S defect annotation regions respectively, and the fine classification hyperparameters indicated by the decision tree. Further, the server can determine the third detection results of the N target images respectively according to the defect region areas corresponding to the S defect annotation regions respectively and the segmentation area hyperparameters indicated by the decision tree. Further, the server can perform multi-perspective decision analysis on the N target images based on the first detection results of the N target images respectively, the second detection results of the N target images respectively, and the third detection results of the N target images respectively, to obtain the image detection results of the N target images respectively.

[0246] For ease of understanding, please refer to Figure 13 , Figure 13 which is a schematic diagram of a scenario for defect quality inspection provided by an embodiment of the present application. It can be understood that the defect types of MIN parts can be mainly divided into two categories, that is, the defect types corresponding to OK defects and the defect types corresponding to NG defects. As Figure 13 shows, it is a schematic diagram of 5 typical defect types in the two types of defect types.

[0247] As Figure 13 shown, the defect schematic diagram corresponding to the defect 13a can be a crack, and the defect 13a can be an NG defect; the defect schematic diagram corresponding to the bright mark can be the defect 13b, and the defect 13b can be an OK defect; the defect schematic diagram corresponding to the missing material can be the defect 13c, and the defect 13c can be an NG defect; the defect schematic diagram corresponding to the bruise can be the defect 13d, and the defect 13d can be an NG defect; the defect schematic diagram corresponding to the dirt can be the defect 13e, and the defect 13e can be an OK defect.

[0248] Among them, it should be understood that when the instance segmentation model and the fine classification model classify the defect annotation areas, the defects that may be either NG defects or OK defects (for example, "bruise") will be considered as NG defects. Furthermore, when performing multi-perspective decision analysis on these defects, "bruise" will be classified as a defect or a non-defect.

[0249] Step S1036: Determine the object detection result of the target object according to the image detection results of the N target images respectively.

[0250] It can be understood that the server can judge the severity of the S defect annotation areas associated with the N target images according to the defect determination rule, obtain the defect severity levels of the S defect annotation areas, and then preferentially output the most serious defect in the target object.

[0251] It should be understood that when the server determines the defect detection result corresponding to the defect annotation area based on the first detection result, the second detection result, and the third detection result, the judgment processes of the first detection result and the second detection result will affect the defect detection result corresponding to the defect annotation area. For easy understanding, please refer to Figure 14 , Figure 14 which is a schematic diagram of a scenario for multi-model comparison provided by an embodiment of the present application. As Figure 14 shown is the Pareto optimization curve of different judgment processes on the validation set, that is, using the Figure 12 shown Pareto optimal method to search for the hyperparameters of the model on the validation set, and then evaluate the results of the model on the test set. Among them, the abscissa can represent the overkill rate, and the ordinate can represent the miss rate.

[0252] It can be understood that after obtaining the overkill and miss rate performance curve of the model using the Pareto optimal algorithm, the appropriate miss and overkill rate can be selected in combination with the business requirements, and the corresponding model hyperparameters (i.e., the target decision hyperparameters) can be selected accordingly. In addition, the Pareto optimal curves of different models also reflect the performance differences of model quality inspection to a certain extent.

[0253] As Figure 14The Pareto optimization curves shown include those corresponding to four types of models: the Pareto optimization curve corresponding to "instance segmentation + fine classification 2" (i.e., pipeline B), the Pareto optimization curve corresponding to "instance segmentation" (i.e., pipeline A), the Pareto optimization curve corresponding to "instance segmentation + fine classification 2" (i.e., pipeline E), and the Pareto optimization curve corresponding to "fine classification 2 + instance segmentation" (i.e., pipeline D). Among them, "instance segmentation" and "fine classification 2" can be seen in Table 2 below.

[0254] Among them, in the quality inspection of a certain MIM part, the offline test data on which the embodiments of the present application are based may include 1324 NG products and 460 OK products. Among them, each NG product and each OK product correspond to multiple images (e.g., N images) at different visual angles. Among them, by controlling overkill to check for missed detections (i.e., the value of the missed detection rate when the overkill rate is 30.435%), and controlling the way of checking for overkill when there is a missed detection (i.e., the value of the overkill rate when the missed detection rate is 2.568%), Figure 14 The analysis results of the Pareto optimization curves shown can be seen in Table 2 below:

[0255] Table 2

[0256]

[0257] Among them, instance segmentation (i.e., maskrcnn(1202det)) means only using the instance segmentation model of maskrcnn for classification (i.e., the first classification information); instance segmentation + fine classification 2 (i.e., maskrcnn(0113cls)) means turning off the instance segmentation model classification and enabling the fine classification model 2 for classification (i.e., the second classification information); instance segmentation + fine classification 1 (i.e., maskrcnn(1202det-0108cls)) means first classifying through the instance segmentation model and then sending the uncertain instances into the fine classification model 1 for classification (i.e., first using the first classification information and then using the second classification information); fine classification 2 + instance segmentation (i.e., maskrcnn(0113cls-1202det)) means first classifying through the fine classification model 2 and then through the instance segmentation model (i.e., first using the second classification information and then using the first classification information); instance segmentation + fine classification 2 (i.e., maskrcnn(1202det-0113cls)) means first classifying through the instance segmentation model and then sending the uncertain instances into the fine classification model 2 for classification (i.e., first using the first classification information and then using the second classification information).

[0258] Among them, compared with the fine classification model 1 (0108clsA), the fine classification model 2 (0113cls) adds image blocks with overkill at fixed positions to the model training, effectively reducing the overkill rate of the system. It can be understood that overkill at fixed positions can be used to eliminate batch defects caused by objective factors. For example, overkill at fixed positions can eliminate batch defects caused by molds. When there are defects in the mold, all target objects generated by the mold have the defects in the mold.

[0259] Among them, it can be seen from Table 2 the impact of the fine classification model on the end-to-end metrics:

[0260] (1) Without classification model: Comparing pipeline A and pipeline C, regardless of the metrics of the fixed overkill rate (30.435%) or the fixed miss rate (2.568%), the miss rate corresponding to pipeline A (the miss rate of pipeline A is 4.607% and the miss rate of pipeline C is 2.266%) and the overkill rate (the overkill rate of pipeline A is 40.652% and the overkill rate of pipeline C is 29.348%) are both higher. This indicates that the fine classification model has a significant positive effect on the end-to-end metrics;

[0261] (2) Updating the classification model: According to the overkill rate analysis results, the fine classification model is updated by collecting image blocks with overkill at fixed positions. Comparing pipeline C and pipeline D, the end-to-end metrics will be significantly improved (the miss rate is reduced from 2.266% to 1.964%, and the overkill rate is reduced from 29.348% to 27.391%). This indicates that collecting image blocks with overkill at fixed positions has a significant positive effect on the end-to-end metrics.

[0262] Among them, it can also be seen from Table 2 the impact of the rule strategy on the end-to-end metrics:

[0263] (1) Comparing pipeline A, pipeline B and pipeline D, pipeline E, if the instance-level defect types completely rely on detection (i.e., instance classification) (pipeline A) or classification (i.e., fine classification) (pipeline B), it is significantly worse than the mode of the fusion of detection and classification (pipeline D and pipeline E), and only using classification (pipeline A) has a worse result than detection (pipeline B);

[0264] (2) Comparing pipeline D and pipeline E, regardless of whether the instance segmentation model classification is used first logically or the fine classification model classification is used first, the Pareto performance curves of the two are close, and the sub-indicators are also relatively close. For example, when controlling the overkill rate to 30.435%, the corresponding miss rates are 1.964% (pipeline D) and 1.813% (pipeline E) respectively; when controlling the miss rate to 2.568%, the corresponding overkill rates are 27.391% (pipeline D) and 25.652% (pipeline E) respectively. Therefore, the effect of pipeline E is better than that of pipeline D.

[0265] It can be understood that the above experiments fully prove the rationality and effectiveness of the quality inspection method proposed in the embodiments of the present application, which first uses an instance segmentation model to highly detect NG defects, then uses a fine-grained classification model to reduce false OK and true NG defects, and then combines industry knowledge for multi-perspective joint inference for quality inspection.

[0266] It should be understood that the factory needs manpower to rejudge the over-killed products. Therefore, the over-kill rate is directly related to the release rate of the production line manpower, and the missed detection rate represents the product quality provided by the factory to the supplier. Generally, under the condition of strictly controlling the missed detection rate, the over-kill rate is minimized to ensure that the product meets the delivery quality and maximize the release of the production line quality inspection manpower. Among them, when rejudging the over-killed products, all the detected products need to be rejudged. Here, all the products can include those that are actually normal and those that are actually defective. For example, the number of defective products among all the products can be 30, the number of actually defective products can be 28, and the number of actually normal products can be 2. These 2 products are the over-killed products.

[0267] It can be seen that the decision analysis model in the embodiments of the present application can combine industry business knowledge and multi-perspective joint inference to make decisions on the object detection results of products (i.e., target objects). Among them, the object detection results of products can be determined through the defect detection results of defects, and the object detection results can determine whether the target object belongs to an NG product or an OK product. Therefore, the embodiments of the present application can achieve the accuracy of quality inspection while realizing the quality inspection of product-level defects and normality, thereby improving the efficiency of quality inspection.

[0268] Further, please refer to Figure 15 , Figure 15 FIG. is a schematic structural diagram of an image data processing device provided by an embodiment of the present application. The image data processing device 1 may include: a first output module 11, a second output module 12, and a decision analysis module 13;

[0269] The first output module 11 is configured to obtain S defect annotation regions associated with N target images, and first defect output results respectively corresponding to the S defect annotation regions; the N target images are obtained by N shooting components respectively shooting the same target object; the visual angles of the N target images are different from each other; N is a positive integer; S is a positive integer; the N target images include a target image Li, where i is a positive integer less than or equal to N;

[0270] Among them, the first output module 11 includes: an image acquisition unit 111 and an instance segmentation unit 112; optionally, the first output module 11 may further include: a label acquisition unit 113, a model output unit 114, and a model training unit 115;

[0271] An image acquisition unit 111 is configured to acquire N target images associated with a target object, and input the N target images into an instance segmentation model respectively;

[0272] An instance segmentation unit 112 is configured to perform instance segmentation on the N target images through the instance segmentation model, to obtain S defect annotation regions associated with the N target images, and first defect output results respectively corresponding to the S defect annotation regions.

[0273] Wherein, the instance segmentation model includes a feature extraction sub-network, a region prediction sub-network and a defect recognition sub-network; the S defect annotation regions include M defect annotation regions in the target image L i ; M is a positive integer less than or equal to S;

[0274] The instance segmentation unit 112 includes: a feature extraction sub-unit 1121, a region prediction sub-unit 1122, and a defect recognition sub-unit 1123;

[0275] The feature extraction sub-unit 1121 is configured to input the target image L i into the feature extraction sub-network, and perform feature extraction on the target image L i through the feature extraction sub-network, to obtain multi-resolution features corresponding to the target image L i ;

[0276] The region prediction sub-unit 1122 is configured to input the multi-resolution features corresponding to the target image L i into the region prediction sub-network, and perform region prediction on the multi-resolution features corresponding to the target image L i through the region prediction sub-network, to obtain M to-be-predicted object regions in the target image L i ;

[0277] The defect recognition sub-unit 1123 is configured to input the M to-be-predicted object regions and the multi-resolution features corresponding to the target image L i into the defect recognition sub-network, and perform defect recognition on the M to-be-predicted object regions and the multi-resolution features corresponding to the target image L i through the defect recognition sub-network, to obtain instance segmentation results respectively corresponding to the M defect annotation regions, first classification probabilities respectively corresponding to the M defect annotation regions, and first classification information respectively corresponding to the M defect annotation regions;

[0278] The defect recognition sub-unit 1123 is configured to use the instance segmentation results respectively corresponding to the M defect annotation regions, the first classification probabilities respectively corresponding to the M defect annotation regions, and the first classification information respectively corresponding to the M defect annotation regions, as the first defect output results respectively corresponding to the M defect annotation regions.

[0279] Among them, the defect recognition subunit 1123 is specifically configured to map M regions of objects to be predicted to the target image L through a defect recognition sub-network i to obtain candidate region features corresponding to the M regions of objects to be predicted respectively;

[0280] The defect recognition subunit 1123 is specifically configured to perform feature alignment on the M candidate region features to obtain aligned region features corresponding to the M candidate region features respectively;

[0281] The defect recognition subunit 1123 is specifically configured to perform a convolution operation on the M aligned region features to obtain classification region features corresponding to the M aligned region features respectively and segmentation region features corresponding to the M aligned region features respectively;

[0282] The defect recognition subunit 1123 is specifically configured to perform a full connection operation on the M classification region features to determine region features corresponding to the M aligned region features respectively and classification features corresponding to the M aligned region features respectively. Based on the M region features, M defect annotation regions are determined. Based on the M classification features, a first classification probability corresponding to each of the M defect annotation regions and a first classification information corresponding to each of the M defect annotation regions are determined;

[0283] The defect recognition subunit 1123 is specifically configured to perform a convolution operation on the M segmentation region features to determine segmentation features corresponding to the M aligned region features respectively. Based on the M segmentation features, an instance segmentation result corresponding to each of the M defect annotation regions is determined.

[0284] Among them, for the specific implementation manners of the feature extraction subunit 1121, the region prediction subunit 1122, and the defect recognition subunit 1123, reference may be made to the descriptions of steps S1012 - S1015 in the corresponding embodiments above Figure 7 which will not be elaborated here.

[0285] Optionally, the label acquisition unit 113 is configured to acquire a defect sample annotation region, defect sample classification information, and a sample boundary region associated with the defect sample image;

[0286] The model output unit 114 is configured to determine a predicted defect annotation region associated with the defect sample image and a first predicted output result corresponding to the predicted defect annotation region in the initial instance segmentation model;

[0287] The model training unit 115 is configured to determine an instance segmentation loss value of the initial instance segmentation model according to the defect sample annotation region, defect sample classification information, sample boundary region, predicted defect annotation region, and the first predicted output result;

[0288] The model training unit 115 is configured to adjust the model parameters in the initial instance segmentation model according to the instance segmentation loss value, and when the adjusted initial instance segmentation model meets the model convergence condition, determine the adjusted initial instance segmentation model as the instance segmentation model.

[0289] Among them, for the specific implementation manners of the image acquisition unit 111, the instance segmentation unit 112, the label acquisition unit 113, the model output unit 114, and the model training unit 115, reference may be made to the descriptions of step S101 and Figure 3 the corresponding embodiment, and Figure 7 the descriptions of steps S1011 - S1015 in the corresponding embodiment, which will not be elaborated here.

[0290] The second output module 12 is configured to determine a second defect output result corresponding to the defect annotation area of the target image L i according to the defect annotation area of the target image L i and the image attribute information of the target image L i ;

[0291] Among them, the image attribute information of the target image L i includes the image serial number of the target image L i and the image output feature corresponding to the target image L i ;

[0292] The second output module 12 includes: a first determination unit 121, a second determination unit 122;

[0293] The first determination unit 121 is configured to determine a defect output feature corresponding to the defect annotation area of the target image L i according to the defect annotation area of the target image L i and the image serial number of the target image L i ;

[0294] Among them, the first determination unit 121 includes: a first determination subunit 1211, a second determination subunit 1212; Optionally, the first determination unit 121 may further include: a label acquisition subunit 1213, a model output subunit 1214, a model training subunit 1215;

[0295] The first determination subunit 1211 is configured to determine the regional coordinates of the defect annotation area of the target image L i , generate a defect input feature corresponding to the defect annotation area of the target image L i according to the regional coordinates and the image serial number of the target image L i , and input the defect input feature into a fine classification model; the fine classification model includes a perceptron sub - network;

[0296] The second determination subunit 1212 is configured to perform a fully connected operation on the defective input features through the sensor subnet to determine the target image L i and the defective output features corresponding to the defective annotation region of

[0297] Optionally, the label acquisition subunit 1213 is configured to acquire the defective sample annotation region and the defective sample classification information associated with the defective sample image, and acquire the normal sample annotation region and the normal sample classification information associated with the normal sample image;

[0298] The model output subunit 1214 is configured to, in the initial fine classification model, determine the second prediction output result corresponding to the defective sample annotation region according to the defective sample annotation region and the image attribute information of the defective sample image, and determine the first classification loss value of the initial fine classification model according to the second prediction output result corresponding to the defective sample annotation region and the defective sample classification information;

[0299] The model output subunit 1214 is configured to determine the second prediction output result corresponding to the normal sample annotation region according to the normal sample annotation region and the image attribute information of the normal sample image, and determine the second classification loss value of the initial fine classification model according to the second prediction output result corresponding to the normal sample annotation region and the normal sample classification information;

[0300] The model training subunit 1215 is configured to determine the fine classification loss value of the initial fine classification model according to the first classification loss value and the second classification loss value;

[0301] The model training subunit 1215 is configured to adjust the model parameters in the initial fine classification model according to the fine classification loss value, and when the adjusted initial fine classification model meets the model convergence condition, determine the adjusted initial fine classification model as the fine classification model.

[0302] Wherein, for the specific implementation manners of the first determination subunit 1211, the second determination subunit 1212, the label acquisition subunit 1213, the model output subunit 1214, and the model training subunit 1215, reference may be made to the descriptions of steps S1021 - S1022 in the corresponding embodiments above, which will not be elaborated here. Figure 9 The second determination unit 122 is configured to determine the second defective output result corresponding to the defective annotation region of the target image L according to the defective output features corresponding to the defective annotation region of the target image L and the image output features corresponding to the target image L

[0303] The second determination unit 122 is configured to determine the second defective output result corresponding to the defective annotation region of the target image L according to the defective output features corresponding to the defective annotation region of the target image L and the image output features corresponding to the target image L i and the defective output features corresponding to the defective annotation region of i and the image output features corresponding to i the target image L

[0304] Wherein, the fine classification model further includes a feature recognition subnet;

[0305] The second determination unit 122 includes: a feature recognition subunit 1221, a feature fusion subunit 1222, and a region classification subunit 1223;

[0306] The feature recognition subunit 1221 is configured to input the target image L i into the feature recognition sub-network, and perform feature recognition on the target image L through the feature recognition sub-network i to obtain the image output features corresponding to the target image L i ;

[0307] The feature fusion subunit 1222 is configured to perform feature fusion on the defect output features corresponding to the defect annotation region of the target image L i and the image output features corresponding to the target image L i to obtain the fusion output features corresponding to the defect annotation region of the target image L i ;

[0308] The region classification subunit 1223 is configured to determine the second defect output result corresponding to the defect annotation region of the target image L based on the fusion output features corresponding to the defect annotation region of the target image L i and the classifier of the fine classification model. i ;

[0309] Specifically, the region classification subunit 1223 is configured to input the fusion output features corresponding to the defect annotation region of the target image L i into the classifier of the fine classification model, and determine the matching degree between the fusion output features corresponding to the defect annotation region of the target image L i and the sample output features in the classifier; the matching degree is used to describe the probability that the defect annotation region of the target image L i belongs to the sample classification label corresponding to the sample output features;

[0310] Specifically, the region classification subunit 1223 is configured to use the sample classification label corresponding to the sample output features with the maximum matching degree as the second classification information corresponding to the defect annotation region of the target image L i , and use the maximum matching degree as the second classification probability corresponding to the defect annotation region of the target image L i ;

[0311] Specifically, the region classification subunit 1223 is configured to use the second classification information corresponding to the defect annotation region of the target image L i and the second classification probability corresponding to the defect annotation region of the target image L i as the second defect output result corresponding to the defect annotation region of the target image L i .

[0312] Among them, for the specific implementation manners of the feature recognition subunit 1221, the feature fusion subunit 1222, and the region classification subunit 1223, reference may be made to the descriptions of steps S1023 - S1025 in the corresponding embodiments above, which will not be elaborated here. Figure 9 For the description of steps S1023 - S1025 in the corresponding embodiments, which will not be elaborated here.

[0313] Among them, for the specific implementation manners of the first determination unit 121 and the second determination unit 122, reference may be made to the descriptions of steps S102 and Figure 3 the corresponding embodiments and Figure 9 For the descriptions of steps S1021 - S1025 in the corresponding embodiments, which will not be elaborated here.

[0314] The decision - making analysis module 13 is configured to perform multi - perspective decision - making analysis on the target object based on the first defect output results respectively corresponding to S defect annotation regions and the second defect output results respectively corresponding to the S defect annotation regions, so as to obtain the object detection result of the target object.

[0315] Among them, the decision - making analysis module 13 includes: a decision tree generation unit 131, a decision - making analysis unit 132, and a result determination unit 133;

[0316] The decision tree generation unit 131 is configured to obtain the business knowledge for performing multi - perspective decision - making analysis on the target object and the target decision hyperparameters associated with the business knowledge, and generate a decision tree according to the business knowledge and the target decision hyperparameters.

[0317] Among them, the decision tree generation unit 131 includes: a set generation subunit 1311 and a decision tree generation subunit 1312;

[0318] The set generation subunit 1311 is configured to obtain the business knowledge for performing multi - perspective decision - making analysis on the target object and a hyperparameter search model, and generate a hyperparameter set associated with the business knowledge through the hyperparameter search model; the hyperparameter set includes one or more groups of decision hyperparameters; each group of decision hyperparameters in the one or more groups of decision hyperparameters includes one or more hyperparameters; the one or more groups of decision hyperparameters are used to balance at least two evaluation indexes corresponding to the decision - making analysis model.

[0319] The decision tree generation subunit 1312 is configured to obtain the target decision hyperparameters that meet the hyperparameter acquisition conditions from the hyperparameter set, and generate a decision tree according to the business knowledge and the target decision hyperparameters.

[0320] Among them, for the specific implementation manners of the set generation subunit 1311 and the decision tree generation subunit 1312, reference may be made to the descriptions of steps S1031 - S1032 in the corresponding embodiments, which will not be elaborated here. Figure 11 For the description of steps S1031 - S1032 in the corresponding embodiments, which will not be elaborated here.

[0321] A decision analysis unit 132, configured to perform multi-perspective decision analysis on N target images in a decision analysis model based on the first defect output results respectively corresponding to S defect annotation regions, the second defect output results respectively corresponding to the S defect annotation regions, and a decision tree, so as to obtain the image detection results of the N target images respectively;

[0322] Among them, the target decision hyperparameters include instance segmentation hyperparameters, segmentation area hyperparameters, and fine classification hyperparameters;

[0323] The decision analysis unit 132 includes: a parameter acquisition unit 1321 and a decision analysis subunit 1322;

[0324] The parameter acquisition unit 1321 is configured to obtain the instance segmentation results respectively corresponding to the S defect annotation regions from the first defect output results respectively corresponding to the S defect annotation regions, and determine the defect region areas respectively corresponding to the S defect annotation regions according to the instance segmentation results respectively corresponding to the S defect annotation regions;

[0325] The parameter acquisition unit 1321 is configured to obtain the first classification probabilities respectively corresponding to the S defect annotation regions and the first classification information respectively corresponding to the S defect annotation regions from the first defect output results respectively corresponding to the S defect annotation regions, and obtain the second classification probabilities respectively corresponding to the S defect annotation regions and the second classification information respectively corresponding to the S defect annotation regions from the second defect output results respectively corresponding to the S defect annotation regions;

[0326] The decision analysis subunit 1322 is configured to perform multi-perspective decision analysis on N target images in the decision analysis model according to the first classification information respectively corresponding to the S defect annotation regions, the second classification information respectively corresponding to the S defect annotation regions, the first classification probabilities respectively corresponding to the S defect annotation regions, the second classification probabilities respectively corresponding to the S defect annotation regions, the defect region areas respectively corresponding to the S defect annotation regions, the S defect annotation regions, and the instance segmentation hyperparameters, segmentation area hyperparameters, and fine classification hyperparameters indicated by the decision tree, so as to obtain the image detection results of the N target images respectively.

[0327] Among them, for the specific implementation manners of the parameter acquisition unit 1321 and the decision analysis subunit 1322, reference can be made to the descriptions of steps S1033 - S1035 in the corresponding embodiments above, which will not be elaborated here. Figure 11 The description of steps S1033 - S1035 in the corresponding embodiments above, which will not be elaborated here.

[0328] A result determination unit 133, configured to determine the object detection result of the target object according to the image detection results of the N target images respectively.

[0329] Among them, for the specific implementation manners of the decision tree generation unit 131, the decision analysis unit 132, and the result determination unit 133, reference may be made to the above Figure 3 corresponding embodiments' descriptions of step S103 and Figure 11 the corresponding embodiments' descriptions of steps S1031 - S1036, which will not be elaborated herein. Additionally, the descriptions of the beneficial effects of using the same method will not be elaborated either.

[0330] Furthermore, please refer to Figure 16 , Figure 16 which is a schematic structural diagram of a computer device provided by an embodiment of the present application. As Figure 16 shown, the computer device 1000 may include: a processor 1001, a network interface 1004, and a memory 1005. In addition, the above computer device 1000 may further include: a user interface 1003 and at least one communication bus 1002. Among them, the communication bus 1002 is used to implement connection communication between these components. Among them, the user interface 1003 may include a display screen (Display) and a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. Optionally, the network interface 1004 may include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a non-volatile memory, for example, at least one disk memory. Optionally, the memory 1005 may further be at least one storage device located far from the aforementioned processor 1001. As Figure 16 shown, the memory 1005, as a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a device control application program.

[0331] In the computer device 1000 as Figure 16 shown, the network interface 1004 can provide network communication functions; while the user interface 1003 is mainly used to provide an input interface for users; and the processor 1001 can be used to call the device control application program stored in the memory 1005 to implement:

[0332] obtaining S defect annotation regions associated with N target images, and first defect output results respectively corresponding to the S defect annotation regions; the N target images are obtained by N shooting components respectively shooting the same target object; the visual angles of the N target images are different from each other; N is a positive integer; S is a positive integer; the N target images include target image L i , where i is a positive integer less than or equal to N;

[0333] According to target image L iThe defect annotation area and the target image L i Based on the image attribute information of, determine the target image L i The second defect output result corresponding to the defect annotation area of

[0334] Based on the first defect output results respectively corresponding to the S defect annotation areas and the second defect output results respectively corresponding to the S defect annotation areas, perform multi-perspective decision analysis on the target object to obtain the object detection result of the target object.

[0335] It should be understood that the computer device 1000 described in the embodiments of the present application can execute the foregoing Figure 3 , Figure 7 , Figure 9 or Figure 11 The description of the image data processing method in the corresponding embodiments, and can also execute the foregoing Figure 15 The description of the image data processing device 1 in the corresponding embodiments will not be elaborated herein. In addition, the description of the beneficial effects of using the same method will not be elaborated either.

[0336] In addition, it should be pointed out here that: the embodiments of the present application also provide a computer-readable storage medium, and the computer-readable storage medium stores the computer program executed by the foregoing image data processing device 1, and the computer program includes program instructions. When the processor executes the program instructions, it can execute the foregoing Figure 3 , Figure 7 , Figure 9 or Figure 11 The description of the image data processing method in the corresponding embodiments, so it will not be elaborated here. In addition, the description of the beneficial effects of using the same method will not be elaborated either. For the technical details not disclosed in the embodiments of the computer-readable storage medium involved in the present application, please refer to the description of the method embodiments of the present application.

[0337] In addition, it should be noted that: the embodiments of the present application also provide a computer program product or a computer program. The computer program product or the computer program may include computer instructions, and the computer instructions may be stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor can execute the computer instructions, so that the computer device executes the foregoing Figure 3 , Figure 7 , Figure 9 or Figure 11 The description of the image data processing method in the corresponding embodiments, so it will not be elaborated here. In addition, the description of the beneficial effects of using the same method will not be elaborated either. For the technical details not disclosed in the embodiments of the computer program product or the computer program involved in the present application, please refer to the description of the method embodiments of the present application.

[0338] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0339] The above-disclosed are only the preferred embodiments of the present application. Of course, the scope of rights of the present application cannot be limited thereby. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.

Claims

1. An image data processing method, characterized in that, Including: Obtaining S defect annotation regions associated with N target images, and first defect output results respectively corresponding to the S defect annotation regions; The N target images are obtained by N shooting components respectively shooting the same target object; visual angles of the N target images are different from each other; N is a positive integer; The S is a positive integer; the N target images include the target image L i , where i is a positive integer less than or equal to N; According to the target image L i 's defect annotation area and the target image L i 's image attribute information, determine the second defect output result corresponding to the defect annotation area of the target image L i ; Obtaining business knowledge for multi-view decision analysis of the target object and target decision hyperparameters associated with the business knowledge, and generating a decision tree according to the business knowledge and the target decision hyperparameters; According to the instance segmentation results in the S first defect output results, determining the defect region areas respectively corresponding to the S defect annotation regions, obtaining the first classification probabilities respectively corresponding to the S defect annotation regions and the first classification information respectively corresponding to the S defect annotation regions from the S first defect output results, and obtaining the second classification probabilities respectively corresponding to the S defect annotation regions and the second classification information respectively corresponding to the S defect annotation regions from the S second defect output results; In a decision analysis model, performing multi-view decision analysis on the N target images according to the first classification information, second classification information, first classification probabilities, second classification probabilities, defect region areas respectively corresponding to the S defect annotation regions, and the target decision hyperparameters indicated by the S defect annotation regions and the decision tree, to obtain image detection results respectively corresponding to the N target images; Determining an object detection result of the target object according to the image detection results respectively corresponding to the N target images.

2. The method according to claim 1, characterized in that, The obtaining S defect annotation regions associated with N target images, and first defect output results respectively corresponding to the S defect annotation regions includes: Obtaining N target images associated with a target object, and respectively inputting the N target images into an instance segmentation model; Performing instance segmentation on the N target images through the instance segmentation model to obtain S defect annotation regions associated with the N target images, and first defect output results respectively corresponding to the S defect annotation regions.

3. The method according to claim 2, wherein The instance segmentation model includes a feature extraction sub-network, a region prediction sub-network, and a defect recognition sub-network; the S defect annotation regions include M defect annotation regions in the target image L i ; the M is a positive integer less than or equal to the S; The performing instance segmentation on the N target images through the instance segmentation model to obtain S defect annotation regions associated with the N target images, and first defect output results respectively corresponding to the S defect annotation regions includes: Input the target image L i into the feature extraction sub-network, and use the feature extraction sub-network to process the target image L i for feature extraction to obtain the multi-resolution features corresponding to the target image L i ; Input the multi-resolution features corresponding to the target image L i into the region prediction sub-network, and perform region prediction on the multi-resolution features corresponding to the target image L i through the region prediction sub-network to obtain M to-be-predicted object regions in the target image L i ; Input the M object regions to be predicted and the multi-resolution features corresponding to the target image L i into the defect recognition sub-network, and perform defect recognition on the multi-resolution features corresponding to the M object regions to be predicted and the target image L i to obtain the instance segmentation results corresponding to the M defect annotation regions, the first classification probabilities corresponding to the M defect annotation regions, and the first classification information corresponding to the M defect annotation regions; Taking the instance segmentation results respectively corresponding to the M defect annotation regions, the first classification probabilities respectively corresponding to the M defect annotation regions, and the first classification information respectively corresponding to the M defect annotation regions as the first defect output results respectively corresponding to the M defect annotation regions.

4. The method according to claim 3, wherein Performing defect recognition on the M object regions to be predicted and the corresponding multi-resolution features of the target image L through the defect recognition sub-network, to obtain the instance segmentation results corresponding to the M defect annotation regions respectively, the first classification probabilities corresponding to the M defect annotation regions respectively, and the first classification information corresponding to the M defect annotation regions respectively, including: i Performing defect recognition on the M object regions to be predicted and the corresponding multi-resolution features of the target image L through the defect recognition sub-network, to obtain the instance segmentation results corresponding to the M defect annotation regions respectively, the first classification probabilities corresponding to the M defect annotation regions respectively, and the first classification information corresponding to the M defect annotation regions respectively, including: Mapping the M to-be-predicted object regions to the target image L through the defect recognition sub-network i to the corresponding multi-resolution features, to obtain candidate region features corresponding to the M to-be-predicted object regions respectively; Performing feature alignment on M candidate region features to obtain aligned region features respectively corresponding to the M candidate region features; Performing convolution operations on the M aligned region features to obtain classification region features respectively corresponding to the M aligned region features and segmentation region features respectively corresponding to the M aligned region features; Perform a fully connected operation on the M classification region features to determine the region features corresponding to the M alignment region features and the classification features corresponding to the M alignment region features respectively. Based on the M region features, determine the M defect annotation regions. Based on the M classification features, determine the first classification probabilities corresponding to the M defect annotation regions and the first classification information corresponding to the M defect annotation regions respectively; Perform a convolution operation on the M segmentation region features to determine the segmentation features corresponding to the M alignment region features respectively. Based on the M segmentation features, determine the instance segmentation results corresponding to the M defect annotation regions respectively.

5. The method according to claim 1, wherein The target image L i 's image attribute information includes the target image L i 's image serial number and the target image L i 's corresponding image output feature; Based on the defect annotation region of the target image L i and the image attribute information of the target image L i to determine the second defect output result corresponding to the defect annotation region of the target image L i includes: According to the defect annotation area of the target image L i and the image serial number of the target image L i to determine the defect output feature corresponding to the defect annotation area of the target image L i ; Based on the defect annotation region corresponding to the defect output feature of the target image L i and the image output feature corresponding to the target image L i determine the second defect output result corresponding to the defect annotation region of the target image L i ​ 6. The method according to claim 5, characterized in that, Based on the defect annotation area of the target image L i and the image serial number of the target image L i to determine the defect output feature corresponding to the defect annotation area of the target image L i includes: Determine the regional coordinates of the defect annotation area of the target image L i According to the regional coordinates and the image serial number of the target image L i generate the defect input features corresponding to the defect annotation area of the target image L i Input the defect input features into the fine classification model; the fine classification model includes a perceptron subnet Perform a fully connected operation on the defective input features through the perception machine subnet to determine the defective output features corresponding to the defective annotation area of the target image L i ​ 7. The method according to claim 6, characterized in that, The fine classification model further includes a feature recognition sub-network; The defect output feature corresponding to the defect annotation area of the target image L i and the image output feature corresponding to the target image L i are used to determine the second defect output result corresponding to the defect annotation area of the target image L i , including: Input the target image L i into the feature recognition sub-network, and perform feature recognition on the target image L through the feature recognition sub-network i to obtain the image output features corresponding to the target image L i ; For the target image L i the defect output features corresponding to the defect annotation area and the target image L i the corresponding image output features are subjected to feature fusion to obtain the target image L i the fusion output features corresponding to the defect annotation area; Based on the fusion output features corresponding to the defect annotation region of the target image L i and the classifier of the fine classification model, determine the second defect output result corresponding to the defect annotation region of the target image L i ​ 8. The method according to claim 7, characterized in that, The fusion output feature corresponding to the defect annotation area based on the target image L i and the classifier of the fine classification model are used to determine the second defect output result corresponding to the defect annotation area of the target image L i as follows: Input the fusion output feature corresponding to the defect annotation region of the target image L i into the classifier of the fine-grained classification model, and determine, through the classifier, the matching degree between the fusion output feature corresponding to the defect annotation region of the target image L i and the sample output feature in the classifier; the matching degree is used to describe the probability that the defect annotation region of the target image L i belongs to the sample classification label corresponding to the sample output feature; Output the sample classification label corresponding to the sample output feature with the maximum matching degree as the target image L i The second classification information corresponding to the defect annotation area of i and use the maximum matching degree as the second classification probability corresponding to the defect annotation area of the target image L Take the second classification information corresponding to the defect annotation area of the target image L i and the second classification probability corresponding to the defect annotation area of the target image L i as the second defect output result corresponding to the defect annotation area of the target image L i ​ 9. The method according to claim 1, characterized in that, The obtaining of the business knowledge for performing multi-perspective decision analysis on the target object and the target decision hyperparameters associated with the business knowledge, and generating a decision tree according to the business knowledge and the target decision hyperparameters, includes: Obtain the business knowledge for performing multi-perspective decision analysis on the target object and a hyperparameter search model, and generate a set of hyperparameters associated with the business knowledge through the hyperparameter search model; the set of hyperparameters includes one or more groups of decision hyperparameters; each group of decision hyperparameters in the one or more groups of decision hyperparameters includes one or more hyperparameters; the one or more groups of decision hyperparameters are used to balance at least two evaluation indicators corresponding to the decision analysis model; Obtain the target decision hyperparameters that meet the hyperparameter obtaining conditions from the set of hyperparameters, and generate a decision tree according to the business knowledge and the target decision hyperparameters.

10. The method according to claim 2, characterized in that, The method further includes: Obtain the defect sample annotation region, the defect sample classification information, and the sample boundary region associated with the defect sample image; In the initial instance segmentation model, determine the predicted defect annotation region associated with the defect sample image and the first predicted output result corresponding to the predicted defect annotation region; According to the defect sample annotation region, the defect sample classification information, the sample boundary region, the predicted defect annotation region, and the first predicted output result, determine the instance segmentation loss value of the initial instance segmentation model; According to the instance segmentation loss value, adjust the model parameters in the initial instance segmentation model. When the adjusted initial instance segmentation model meets the model convergence condition, determine the adjusted initial instance segmentation model as the instance segmentation model.

11. The method according to claim 6, characterized in that, The method further includes: Obtain the defect sample annotation region and the defect sample classification information associated with the defect sample image, and obtain the normal sample annotation region and the normal sample classification information associated with the normal sample image; In the initial fine classification model, according to the defect sample annotation region and the image attribute information of the defect sample image, determine the second predicted output result corresponding to the defect sample annotation region. According to the second predicted output result corresponding to the defect sample annotation region and the defect sample classification information, determine the first classification loss value of the initial fine classification model; Determine the second predicted output result corresponding to the normal sample annotation region according to the image attribute information of the normal sample annotation region and the normal sample image, and determine the second classification loss value of the initial fine-grained classification model according to the second predicted output result corresponding to the normal sample annotation region and the normal sample classification information; Determine the fine-grained classification loss value of the initial fine-grained classification model according to the first classification loss value and the second classification loss value; Adjust the model parameters in the initial fine-grained classification model according to the fine-grained classification loss value. When the adjusted initial fine-grained classification model meets the model convergence condition, determine the adjusted initial fine-grained classification model as the fine-grained classification model.

12. A computer device, characterized in that, Including: A processor and a memory; The processor is connected to the memory. Among them, the memory is used to store a computer program, and the processor is used to call the computer program so that the computer device executes the method according to any one of claims 1-11.

13. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and the computer program is suitable for being loaded and executed by a processor so that a computer device having the processor executes the method according to any one of claims 1-11.

14. A computer program product, characterized in that, The computer program product includes computer instructions, and the computer instructions are stored in the computer-readable storage medium and are suitable for being read and executed by a processor so that a computer device having the processor executes the method according to any one of claims 1-11.

Citation Information

Patent Citations

  • Defect detection model training method, defect detection method and related device

    CN111814850A

  • Target detection method and device and electronic equipment

    CN112634201A