Image processing method and device, electronic equipment, storage medium and program product
By using prototype vectors to correct the image features of the student model in multiple cascaded network layers of the teacher and student models, the problem of small feature differences between the teacher and student models in defective image detection is solved, thereby improving the accuracy and generalization ability of defect detection.
Patent Information
- Application Number
- CN202410486729.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-22
- Publication Date
- 2025-10-24
AI Technical Summary
In existing defect detection methods based on knowledge distillation, the differences in features extracted by the teacher model and the student model when processing defective images are small, resulting in low accuracy of defect detection results.
By using multiple cascaded network layers of trained teacher and student models, image features are acquired. A preset number of prototype vectors are used to correct the query image features of the student model, making them merged towards the direction of defect-free objects. This expands the difference in query image features output by the student and teacher models, thereby achieving effective defect detection.
It improves the accuracy of defect detection results, reduces the impact of factors such as shooting angle and lighting conditions on the detection results, and enhances the generalization performance of the model.
Smart Images

Figure CN120833294A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of artificial intelligence, and in particular, the present disclosure relates to an image processing method and device, an electronic device, a storage medium, and a program product. BACKGROUND
[0002] With the development of technology, computer vision technology is increasingly applied in industrial scenarios. For example, in modern industrial manufacturing, mechanical parts, electronic components, and other products produced by industrial assembly lines inevitably have defects. Therefore, in the quality control link, product images can be used for image processing to achieve defect detection of the products.
[0003] Currently, a defect detection method based on knowledge distillation usually performs unsupervised learning on a student model based on a defect-free image (i.e., an image corresponding to a defect-free product) and a trained teacher model, and transfers the knowledge of the teacher model to the student model, thereby helping the student model better learn the capabilities of the teacher model.
[0004] Since the teacher model only teaches the student model the ability to extract features of a defect-free image, the student model does not process images with defects, and thus the teacher model and the student model extract different features when processing images with defects, thereby enabling the use of such differences for defect detection. However, the teacher model and the student model have similar structures, which can easily lead to similar results obtained by the teacher model and the student model when processing images with defects, thereby resulting in ineffective defect detection and low accuracy of the defect detection result. SUMMARY
[0005] The present disclosure provides an image processing method, device, electronic device, storage medium, and program product, which can solve the problem of low accuracy of the defect detection result in the prior art. The technical solutions provided by the present disclosure are as follows:
[0006] According to an aspect of an embodiment of the present disclosure, an image processing method is provided, which includes:
[0007] obtaining a query image for a target category to be detected and a support image corresponding to the query image;
[0008] obtaining first query image features of the query image corresponding to each level respectively and support image features of the support image corresponding to each level respectively through a plurality of cascaded first network layers of a trained teacher model;
[0009] obtaining second query image features of the query image corresponding to each level respectively through a plurality of cascaded second network layers of a trained student model; the second image features represent image features obtained after correcting the query image;
[0010] determining a first target image feature from a plurality of levels of first query image features and determining a second target image feature from a plurality of levels of second query image features; determining a defect detection result for the query image based on a comparison result between the first target image feature and the second target image feature;
[0011] wherein, for each level, the second query image feature is obtained by:
[0012] extracting a preset number of prototype vectors based on the support image feature of the previous level; the preset number of prototype vectors are used to represent information of a defect-free object of a target class;
[0013] fusing the second query image feature of the previous level and the preset number of prototype vectors to obtain a fusion feature, and inputting the fusion feature into a second network layer of the corresponding level to obtain the second query image feature of the level.
[0014] Optionally, the fusing the second query image feature of the previous level and the preset number of prototype vectors to obtain a fusion feature comprises:
[0015] determining a plurality of similarity information between the second query image feature of the previous level and the preset number of prototype vectors, respectively;
[0016] fusing the plurality of similarity information to obtain a first fusion feature;
[0017] taking the prototype vector with the largest similarity to the second query image feature of the previous level as a guide feature;
[0018] fusing the second query image feature of the previous level, the first fusion feature and the guide feature to obtain the fusion feature.
[0019] Optionally, the student model is obtained by training based on:
[0020] performing at least one training operation on the initial student model based on a plurality of sample sets and the trained teacher model until a training end condition is met, and taking the initial student model that meets the training end condition as the trained student model;
[0021] The sample set includes a sample query image for a sample class and at least one sample support image corresponding to the sample query image.
[0022] wherein, the training operation comprises:
[0023] extracting a sample support image feature corresponding to the sample support image through the teacher model, and extracting a first sample query image feature corresponding to the sample query image through the teacher model;
[0024] extracting a preset number of sample prototype vectors corresponding to the sample support image feature based on the sample support image feature;
[0025] fusing the second sample query image feature and the preset number of sample prototype vectors to obtain a sample fusion feature, and determining a total training loss based on the sample fusion feature and the first sample query image feature;
[0026] adjusting parameters of the initial student model based on the total training loss, and taking the initial student model after adjusting the parameters as an initial student model corresponding to a next training operation.
[0027] Optionally, the determining of the total training loss based on the sample fusion feature and the first sample query image feature comprises:
[0028] For each level, determining a corresponding second sample query image feature based on the sample fusion feature through a second network layer of the level;
[0029] For each level, determining a first training loss of the level based on a difference between the first sample query image feature corresponding to the first network layer of the level and the second sample query image feature corresponding to the second network layer of the level;
[0030] determining the total training loss based on the first training losses respectively corresponding to the multiple levels.
[0031] Optionally, for each level, the determining of the first training loss of the level based on the difference between the first sample query image feature corresponding to the first network layer of the level and the second sample query image feature corresponding to the second network layer of the level comprises:
[0032] determining a second training loss based on the difference between the first sample query image feature corresponding to the first network layer of the level and the second sample query image feature corresponding to the second network layer of the level;
[0033] determining a third training loss based on orthogonality between the preset number of sample prototype vectors corresponding to the level;
[0034] determining the first training loss of the level based on the second training loss and the third training loss.
[0035] Optionally, the method further comprises:
[0036] For each sample support image, reference information is determined based on a correlation between a sample support image feature corresponding to the sample support image and a first sample query image feature corresponding to a target sample query image; the target sample query image and the sample support image belong to a same sample set;
[0037] Based on a correlation between the sample support image feature corresponding to the sample support image and a second sample query image feature corresponding to a sample query image in each sample set, to-be-compared information is determined;
[0038] Based on a difference between the reference information and the to-be-compared information, a fourth training loss corresponding to the sample support image is determined;
[0039] Based on the fourth training loss corresponding to each sample support image respectively, a fifth training loss is determined;
[0040] The first training loss is determined based on at least one of the following manners:
[0041] The first training loss is determined based on the second training loss and the fifth training loss;
[0042] The first training loss is determined based on the second training loss, the third training loss, and the fifth training loss.
[0043] Optionally, the to-be-compared information is determined based on a correlation between the sample support image feature corresponding to the sample support image and a second sample query image feature corresponding to a sample query image in each sample set, comprising:
[0044] First to-be-compared information is determined based on a correlation between the sample support image feature corresponding to the sample support image and a second sample query image feature corresponding to a target query image;
[0045] Second to-be-compared information is determined based on a correlation between the sample support image feature corresponding to the sample support image and a second sample query image feature corresponding to a non-target sample query image;
[0046] The fourth training loss corresponding to the sample support image is determined based on a difference between the reference information and the to-be-compared information, comprising:
[0047] The fourth training loss corresponding to the sample support image is determined based on a difference between the reference information and the first to-be-compared information, and a difference between the reference information and the second to-be-compared information.
[0048] According to another aspect of the embodiments of the present disclosure, an image processing apparatus is provided, which comprises:
[0049] an image acquisition module, configured to acquire a query image of a target category to be detected and a support image corresponding to the query image;
[0050] a first feature acquisition module, configured to acquire first query image features of the query image corresponding to each level respectively and support image features of the support image corresponding to each level respectively by using a plurality of cascaded first network layers of the trained teacher model;
[0051] a second feature acquisition module, configured to acquire second query image features of the query image corresponding to each level respectively by using a plurality of cascaded second network layers of the trained student model; the second image features represent image features of the query image after correction;
[0052] a feature comparison module, configured to determine first target image features from the first query image features of the plurality of levels and determine second target image features from the second query image features of the plurality of levels; and determine a defect detection result for the query image based on a comparison result between the first target image features and the second target image features;
[0053] For each level, the second query image features are acquired by the following manner:
[0054] extracting a preset number of prototype vectors based on the support image features of the previous level; the preset number of prototype vectors are used to represent information of a defect-free object of the target category;
[0055] fusing the second query image features of the previous level and the preset number of prototype vectors to obtain fused features, and inputting the fused features into the second network layer of the corresponding level to obtain the second query image features of the level.
[0056] According to another aspect of the embodiments of the present disclosure, an electronic device is provided, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the steps of any of the above image processing methods when executing the program.
[0057] According to still another aspect of the embodiments of the present disclosure, a computer readable storage medium is provided, which stores a computer program, and the computer program is executable on a processor to implement the steps of any of the above image processing methods.
[0058] According to an aspect of the embodiments of the present disclosure, a computer program product is provided, which includes a computer program executable on a processor to implement the steps of any of the above image processing methods.
[0059] The technical scheme provided by the embodiments of the present disclosure has the beneficial effects that:
[0060] The teacher model obtains a preset number of prototype vectors for representing information of the defect-free object based on the support image, and the second query image feature is fused with the preset number of prototype vectors, so that the student model extracts the second query image feature to the direction of the defect-free object by using the prototype vector, effectively suppresses the defect information of the query image extracted by the student model, makes the feature extracted by the student model as close as possible to the feature of the defect-free image corresponding to the query image, expands the difference between the query image features output by the student model and the teacher model, and effectively detects defects, thereby improving the accuracy of the defect detection result. BRIEF DESCRIPTION OF DRAWINGS
[0061] In order to more clearly illustrate the technical schemes in the embodiments of the present disclosure, the drawings needed to be used in the description of the embodiments of the present disclosure will be briefly introduced.
[0062] Figure 1 The application environment schematic diagram of the image processing method provided by the embodiments of the present disclosure is shown in the figure;
[0063] Figure 2 The flowchart of the image processing method provided by the embodiments of the present disclosure is shown in the figure;
[0064] Figure 3 The flowchart of the reverse knowledge distillation provided by the embodiments of the present disclosure is shown in the figure;
[0065] Figure 4 The flowchart of the forward knowledge distillation provided by the embodiments of the present disclosure is shown in the figure;
[0066] Figure 5 The flowchart of the second query image feature determination process provided by the embodiments of the present disclosure is shown in the figure;
[0067] Figure 6 The schematic diagram of the prototype vector extraction and integration process provided by the embodiments of the present disclosure is shown in the figure;
[0068] Figure 7 The flowchart of the model training method provided by the embodiments of the present disclosure is shown in the figure;
[0069] Figure 8 The structural schematic diagram of the image processing device provided by the embodiments of the present disclosure is shown in the figure;
[0070] Figure 9 The structural schematic diagram of the electronic device provided by the embodiments of the present disclosure is shown in the figure. DETAILED DESCRIPTION
[0071] Embodiments of the present disclosure will be described below with reference to the accompanying drawings. It should be understood that the embodiments described below in conjunction with the drawings are exemplary descriptions of the technical solutions of the embodiments of the present disclosure, and do not constitute a limitation on the technical solutions of the embodiments of the present disclosure.
[0072] Those skilled in the art can understand that the singular forms "a", "an" and "the" used herein include plural forms unless specifically stated otherwise. It should be further understood that the terms "comprise" and "include" used in the embodiments of the present disclosure mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements and / or components, but do not exclude other features, information, data, steps, operations, elements, components and / or combinations thereof supported by the present technology. It should be understood that when we say that an element is "connected" or "coupled" to another element, the element can be directly connected or coupled to the other element, or it can mean that the element and the other element are connected through an intermediate element. In addition, "connected" or "coupled" used herein can include wireless connection or wireless coupling. The term "and / or" used herein indicates that at least one of the items defined by the term, for example, "A and / or B" or "A, B" indicates implementation as "A", or implementation as "B", or implementation as "A and B".
[0073] In order to make the purposes, technical solutions and advantages of the present disclosure clearer, the embodiments of the present disclosure will be described in further detail below with reference to the accompanying drawings.
[0074] First, the technical terms related to the present disclosure are introduced and explained:
[0075] Knowledge Distillation: Knowledge distillation is a model compression technique that aims to improve the performance of one model (usually a simpler student model) by having another model (usually a more complex teacher model) guide it. The teacher model and the student model are two models in knowledge distillation, the teacher model guides the student model, and the knowledge of the teacher model is transferred to the student model, so as to help the student model better learn the ability of the teacher model. Among them, the teacher model is a trained model, the model parameters of the teacher model are fixed, the network structure of the teacher model is more complex, the model parameters are more, but the accuracy is higher; the network structure of the student model is simpler, the model parameters are fewer. In the training process, the student model will try to simulate the output of the teacher model, and through the method of knowledge distillation, the knowledge provided by the teacher model is used to optimize the model parameters of the student model and improve the performance of the student model.
[0076] Optionally, the image processing method according to the embodiments of the present disclosure can be implemented based on machine learning (ML) and computer vision (CV) in artificial intelligence (AI).
[0077] Artificial intelligence is a theory, method, technology and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence, perceive an environment, acquire knowledge and use the knowledge to obtain optimal results. In other words, artificial intelligence is a comprehensive technology of computer science, which attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that the machines have the functions of perception, reasoning and decision-making.
[0078] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, both hardware and software technologies. Artificial intelligence basic technologies generally include sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, automatic driving, intelligent transportation and other major directions.
[0079] Machine learning (ML) is a multi-disciplinary subject involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It is a discipline that studies how computers simulate or implement human learning behavior to acquire new knowledge or skills, reorganize existing knowledge structure to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental approach to making computers intelligent, and its applications are widespread in various fields of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rule-based learning.
[0080] Computer Vision (CV) is a science that studies how to make machines "see". More specifically, it refers to using cameras and computers to replace human eyes to identify and measure targets, and further perform image processing to make computer processing more suitable for human eye observation or image transmission to instruments for detection. As a scientific discipline, computer vision researches related theories and technologies, and attempts to establish artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and other technologies. It also includes common face recognition, fingerprint recognition, and other biometric identification technologies.
[0081] Optionally, the data processing involved in the method provided by the embodiments of the present disclosure can also be implemented based on cloud technology. For example, various calculations in model training can be implemented using cloud computing technology, and training data can be stored using cloud storage.
[0082] Among them, cloud computing (cloud computing) is a computing mode that distributes computing tasks on a large number of computing resources, so that various application systems can obtain computing power, storage space and information services according to needs, and the network providing resources is called "cloud". The resources in the "cloud" are infinitely expandable to the user, and can be obtained at any time, used on demand, expanded at any time, and paid according to use. Cloud storage (cloud storage) is a new concept extended and developed on the basis of the concept of cloud computing. Cloud storage can collect a large number of various types of storage devices (storage nodes) in the network through application software or application interface to work together and provide a storage system for data storage and business access functions.
[0083] In the specific embodiments of the present disclosure, any data related to the object is involved, and when the embodiments of the present disclosure are applied to specific products or technologies, the permission or consent of the object is required, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of the country and region. That is, if any of the above data related to the object is involved in the embodiments of the present disclosure, these data need to be obtained with the authorization and consent of the object, and in compliance with the relevant laws, regulations and standards of the country and region.
[0084] The technical solutions of the embodiments of the present disclosure and the technical effects brought by the technical solutions of the present disclosure are described below through the description of several exemplary embodiments. It should be noted that the following embodiments can be mutually referenced, borrowed or combined. For the same terms, similar features and similar implementation steps in different embodiments, they are not described repeatedly.
[0085] Figure 1 An application environment schematic diagram of the image processing method provided by the embodiments of the present disclosure is shown. The application environment can include a server 101 and a terminal 102. The server 101 can obtain a query image for a target category to be detected and a support image corresponding to the query image from the terminal 102. The server 101 obtains first query image features of the query image corresponding to each level respectively by a plurality of cascaded first network layers of a trained teacher model, and obtains support image features of the support image corresponding to each level respectively. The server 101 obtains second query image features of the query image corresponding to each level respectively by a plurality of cascaded second network layers of a trained student model. The server 101 determines a first target image feature from the first query image features of the plurality of levels, and determines a second target image feature from the second query image features of the plurality of levels. The server 101 determines a defect detection result for the query image based on a comparison result between the first target image feature and the second target image feature, and returns the defect detection result to the terminal 102.
[0086] The server 101 can also serve as a training server to train an initial student model based on a plurality of sample sets and a teacher model to obtain a trained student model. An additional training server can also be set up, and the trained teacher model and the student model can be deployed on the server 101.
[0087] The image processing method provided by the embodiments of the present disclosure can be executed by any electronic device. The electronic device can be a server or a terminal as shown. Figure 1
[0088] The server can be a stand-alone physical server, a server cluster formed by multiple physical servers, or a distributed system, and can also be a cloud server or a server cluster providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and basic cloud computing services such as big data and artificial intelligence platforms. The terminal can be a smartphone (such as an Android phone, an iOS phone, etc.), a tablet computer, a notebook computer, a digital broadcast receiver, a MID (Mobile Internet Device), a PDA (Personal Digital Assistant), a desktop computer, a smart home appliance, a vehicle-mounted terminal (such as a vehicle-mounted navigation terminal, a vehicle-mounted computer, etc.), a smart speaker, a smart watch, etc. The terminal and the server can be connected directly or indirectly through wired or wireless communication, but are not limited thereto.
[0089] Figure 2 A flowchart of an image processing method provided by an embodiment of the present disclosure is shown in FIG. 1, which includes the following steps. Figure 2
[0090] In step S110, a query image for a target category to be detected and a support image corresponding to the query image are obtained.
[0091] Specifically, the target category can be a category of a product included in an image, and the query image to be detected can be an image of a product that needs to be detected for defects. The product in the image can be an industrial product, such as an industrial part or a component.
[0092] The support image corresponding to the query image can be an image of a non-defective product belonging to the same target category as the query image. The query image can correspond to one support image, or can correspond to multiple support images.
[0093] The query image and the support image can be obtained by photographing through an image acquisition device (such as a camera, a video camera, a camera, a scanner, a scanning pen, a mobile phone, or a tablet computer, etc.), or can be obtained by collecting through a network under the premise of complying with relevant regulations, which is not limited in the present disclosure. For example, the support image of a non-defective product can be obtained by photographing a part or the whole of the product when the product is shipped.
[0094] In step S120, first query image features of the query image corresponding to each level are obtained through a plurality of cascaded first network layers of a trained teacher model, and support image features of the support image corresponding to each level are obtained.
[0095] Step S130, obtaining second query image features of the query image corresponding to each layer through multiple cascaded second network layers of the trained student model; the second image features represent image features obtained after correcting the query image;
[0096] For each level, the second query image feature is obtained in the following manner:
[0097] Extracting a preset number of prototype vectors based on the supporting image features of the previous layer; the preset number of prototype vectors are used to represent information of defect-free objects of the target category;
[0098] The second query image feature of the previous level is fused with a preset number of prototype vectors to obtain a fused feature, and the fused feature is input into the second network layer of the corresponding level to obtain the second query image feature of the level.
[0099] Specifically, in the embodiment of the present disclosure, the student model is obtained by training the initial student model based on the trained teacher model and multiple sample sets. The training method of the student model will be explained in detail below.
[0100] The trained teacher model includes multiple cascaded first network layers, and the trained student model includes multiple cascaded second network layers. The number of first network layers in the teacher model and the number of second network layers in the student model can be the same.
[0101] Optionally, the teacher model and the student model can be built based on the reverse knowledge distillation paradigm. Figure 3 A schematic diagram of a reverse knowledge distillation process provided by an embodiment of the present disclosure is shown as follows: Figure 3 As shown in the figure, the gray trapezoids represent the first network layer, the white trapezoids represent the second network layer, the teacher model includes four gray cascaded trapezoids a, b, c, and d, and the student model includes four gray cascaded trapezoids e, f, g, and h. During the reverse knowledge distillation process, the query image and support images are input to the teacher model, and the output of the teacher model is then input to the student model.
[0102] In reverse knowledge distillation, the first network layer of the teacher model and the second network layer of the student model undergo data processing in reverse order. For example, the teacher model can include multiple cascaded encoders, and the student model can include multiple cascaded decoders. The student model reconstructs features based on the teacher model's output and compares the reconstructed features output by the student model with the original features extracted by the teacher model, thereby achieving defect detection in the query image.
[0103] Optionally, the teacher model and the student model can be constructed based on the forward knowledge distillation paradigm. Figure 4A flowchart of a forward knowledge distillation provided by an embodiment of the present disclosure is shown in FIG. 1. Figure 4 As shown in FIG. 1, the gray trapezoids represent the first network layers, and the white trapezoids represent the second network layers. The teacher model includes four cascaded gray trapezoids a, b, c, and d, and the student model includes four cascaded gray trapezoids e, f, g, and h. In the process of forward knowledge distillation, the query image and the support image are input into the teacher model, and the query image is input into the student model. In the forward knowledge distillation, the data processing process of the first network layer of the teacher model is the same as that of the second network layer of the student model.
[0104] The query image and the corresponding support image are input into the teacher model. The query image is subjected to multi-scale feature extraction through the multiple cascaded first network layers in the teacher model, and the first query image features corresponding to respective levels are obtained. The support image is subjected to multi-scale feature extraction, and support image features corresponding to respective levels are obtained.
[0105] In the process of forward knowledge distillation, the query image is input into the student model, and the query image is subjected to multi-scale feature extraction through the multiple cascaded second network layers in the student model, and second query image features corresponding to respective levels are obtained. In the process of backward knowledge distillation, the first query image features output by the teacher model are input into the student model, and the first query image features are subjected to feature reconstruction in sequence through the multiple cascaded second network layers in the student model, and second query image features corresponding to respective levels are obtained.
[0106] For each level, the support image features output by the first network layer of the previous level can be obtained, and a preset number of prototype vectors corresponding thereto can be extracted based on the support image features, wherein the preset number of prototype vectors can be used to represent the information of the defect-free object of the target category (i.e., the defect-free product corresponding to the target category).
[0107] It should be noted that the support image in the embodiment of the present disclosure is the image corresponding to the defect-free object, and the preset number of prototype vectors extracted based on the support image features can include the essential information of the defect-free object of the target category. As can be known by those skilled in the art, the preset number of prototype vectors is equivalent to a plurality of basis vectors, i.e., the support image of the defect-free object can be obtained by linear transformation of the preset number of prototype vectors.
[0108] The second query image features output by the second network layer of the previous level are obtained, and the second query image features of the previous level are fused with the preset number of prototype vectors to obtain fusion features. The fusion features are input into the second network layer of the level, and the fusion features are subjected to feature reconstruction by the second network layer to obtain the second query image features corresponding to the level.
[0109] In the embodiments of the present disclosure, the preset number of prototype vectors can represent the information of the defect-free object of the target category. By fusing the second query image feature with the preset number of prototype vectors, the prototype vectors are used to correct the second query image feature extracted by the student model towards the direction of the defect-free object, effectively suppress the defect information of the query image extracted by the student model, so that the feature extracted by the student model is as close as possible to the feature of the defect-free image corresponding to the query image.
[0110] Figure 5 A flowchart of a second query image feature determination process provided by the embodiments of the present disclosure is shown in FIG. 1. As shown in FIG. 1, the gray trapezoid represents the first network layer, and the white trapezoid represents the second network layer. The first network layer output of the i-th level corresponds to the first query image feature of the i-th level Figure 5 The second network layer output of the i-th level corresponds to the second query image feature of the i-th level The support image feature of the i-th level corresponds to the support image feature of the i-th level The second network layer output of the i-th level corresponds to the second query image feature of the i-th level The support image feature of the i-th level corresponds to the support image feature of the i-th level A preset number of prototype vectors are extracted, and the preset number of prototype vectors are fused with the second query image feature of the i-th level to obtain the fusion feature The fusion feature is input into the second network layer of the i+1-th level to obtain the second query image feature output by the second network layer of the i+1-th level
[0111] It should be noted that the "level" in the embodiments of the present disclosure is not a concept on the model structure. The "level" can be understood as a logical concept, that is, the first network layer and the second network layer of the same level can be understood as network layers with consistent output feature sizes, that is, the first network layer or the second network layer of the i+1-th level is not necessarily the network layer after the first network layer or the second network layer of the i-th level in the model. For example, when the first network layer is an encoder and the second network layer is a decoder, the first network layer and the second network layer of the same level are the encoder and the decoder of the same level.
[0112] In step S140, the first target image feature is determined from the first query image features of multiple levels, and the second target image feature is determined from the second query image features of multiple levels. Based on the comparison result between the first target image feature and the second target image feature, the defect detection result for the query image is determined.
[0113] Specifically, the first query image features corresponding to each level can be obtained through the plurality of cascaded first network layers in the trained teacher model; and the second query image features corresponding to each level can be obtained through the plurality of cascaded second network layers in the trained student model.
[0114] The first target image feature can be determined from the first query image features corresponding to each level, and the second target image feature can be determined from the second query image features corresponding to each level.
[0115] The first target image feature and the second target image feature can be of the same feature scale. That is, the first target image feature can be an image feature extracted by any first network layer in the teacher model based on the query image, and the second target image can be an image feature output by a second network layer in the student model at the same level as the first network layer based on the query image. In forward knowledge distillation, the second target image feature can be an image feature extracted by the second network layer based on the query image; and in reverse knowledge distillation, the second target image feature can be a reconstructed feature obtained by the second network layer based on the output of the teacher model.
[0116] Since the teacher model has good feature expression capability, the teacher model can accurately extract the actual image feature of the query image, and the student model extracts a feature as close as possible to the feature of the corresponding defect-free image of the query image. By comparing the first target image feature output by the teacher model and the second target image feature output by the student model, and using the difference between the first target image feature and the second target image, a defect detection result for the query image can be obtained.
[0117] When the first target image feature and the second target image feature are consistent, for example, the similarity between the first target image feature and the second target image feature is greater than a preset threshold, it indicates that the actual image feature of the query image is close to the image feature of the corresponding defect-free image, and the product object included in the query image is more likely to be a defect-free object, and it can be determined that the defect detection result is that the product is qualified.
[0118] When the first target image feature and the second target image feature are inconsistent, for example, the similarity between the first target image feature and the second target image feature is not greater than a preset threshold, it indicates that the actual image feature of the query image is greatly different from the image feature of the corresponding defect-free image, and the product object included in the query image is more likely to be a defective object, and it can be determined that the defect detection result is that the product is unqualified.
[0119] Optionally, when the first target image feature is inconsistent with the second target image feature, a position where the product appears a defect in the query image can also be located based on a difference between the first target image feature and the second target image feature in a spatial dimension, and the located defect position is taken as the defect detection result.
[0120] In the embodiments of the present disclosure, the teacher model obtains a preset number of prototype vectors for representing information representing a defect-free object based on the support image, and the second query image feature is fused with the preset number of prototype vectors, so that the prototype vectors correct the second query image feature extracted by the student model towards the direction of the defect-free object, effectively suppress the defect information of the query image extracted by the student model, so that the feature extracted by the student model is as close as possible to the feature of the defect-free image corresponding to the query image, the difference between the query image features output by the student model and the teacher model is enlarged, and effective defect detection can be realized, and the accuracy of the defect detection result is improved.
[0121] In addition, in the embodiments of the present disclosure, the image features of the query image output by the student model and the teacher model are compared, compared with the prior art of directly comparing the query image with the standard support image, the influence of the inconsistency between the query image and the support image caused by other interference factors (such as shooting angle, lighting condition, background setting, etc.) on the defect detection result is effectively reduced, and the accuracy of the defect detection result is further improved.
[0122] As an optional embodiment, in the method, for each level, the second query image feature of the level is fused with the preset number of prototype vectors to obtain a fused feature, including:
[0123] determining a plurality of similarity information between the second query image feature and the preset number of prototype vectors;
[0124] fusing the plurality of similarity information to obtain a first fused feature;
[0125] taking the prototype vector with the maximum similarity with the second query image feature as a guide feature;
[0126] fusing the second query image feature, the first fused feature and the guide feature to obtain the fused feature.
[0127] Specifically, for each level, after obtaining the preset number of prototype vectors corresponding to the support image feature of the level, the similarity information between the second query image feature output by the second network layer of the previous level and the preset number of prototype vectors can be calculated, and the similarity information can be a cosine similarity.
[0128] The similarity information between the second query image feature of the previous level and the preset number of prototype vectors can be fused to obtain the corresponding first fusion feature. The prototype vector with the maximum similarity with the second query image feature of the previous level can also be taken as the corresponding guide feature, and the second query image feature of the previous level, the first fusion feature and the guide feature are fused to obtain the corresponding fusion feature of the level.
[0129] The guide feature can be understood as prior information. Through the guide feature, the model can learn the features of the data more effectively, and can also help the model to learn a representation with better generalization ability, so that the model has better adaptability to unseen data. By introducing prior information, the model can learn a more general feature representation, thereby improving the generalization performance of the model.
[0130] Figure 6 A schematic diagram of a prototype vector extraction and integration process provided by the embodiments of the present disclosure is shown in Figure 6 It is assumed that represents the support image feature output by the first network layer of the i-th level, where C represents the channel number of the feature, H represents the height of the feature, W represents the width of the feature, and m represents any pixel. A convolution (conv) with a size of 3x3 can be used: and a reshape operation to transform them into the feature L represents a preset number; softmax functions (a kind of activation function) are used along the spatial dimension of to generate L attention maps; each attention map is multiplied element by element with and the resulting result is aggregated along the spatial dimension of to generate L prototype vectors
[0131] The generation process of the prototype vector can be represented by the following formula:
[0132]
[0133] In the formula, (h, w) represents the spatial coordinates.
[0134] After obtaining the L prototype vectors, the cosine similarity between each prototype vector and the second query image feature at position (h, w) of the second network layer output of the i-th level is calculated. The specific calculation formula is as follows:
[0135]
[0136] In the formula, the symbol ":=" represents definition, The similarity between the prototype vector and the query image feature is calculated for each prototype vector in the prototype vector set. For position (h, w), the prototype vector with the maximum similarity is selected as the guide feature wherein argmax l represents the maximum value of the subscript l. The The similarity information is added along the channel direction to obtain a probability map (i.e., the first fused feature). Finally, the second query image feature The guide feature and the probability map will be spliced along the channel direction, and the spliced result is input into the second network layer of the i+1 level to obtain the corresponding second query image feature of the i+1 level output by the second network layer of the i+1 level The second query image feature The specific formula can be as follows:
[0137]
[0138] In the formula, represents a splicing operation, D i+1 represents the second network layer of the i+1 level.
[0139] In the embodiments of the present disclosure, the similarity information between the second query image feature of the previous level and a preset number of prototype vectors is fused to obtain the corresponding first fused feature; the prototype vector with the maximum similarity with the second query image feature of the previous level is taken as the corresponding guide feature, and the second query image feature of the previous level, the first fused feature and the guide feature are fused to obtain the fusion feature of the level, so that the fusion feature can not only include the information of the prototype vector closest to the query image, but also include the similarity information between the query image and multiple prototype vectors. The second query image feature extracted by the student model is corrected in the direction of the defect-free object by using different dimensional information, the defect information of the query image extracted by the student model is effectively suppressed, and the feature extracted by the student model is as close as possible to the feature of the defect-free image corresponding to the query image.
[0140] As an optional embodiment, the student model is trained based on the following manner:
[0141] The initial student model is trained at least once based on the multiple sample sets and the trained teacher model until the training end condition is met, and the initial student model meeting the training end condition is taken as the trained student model.
[0142] The sample set includes a sample query image for a sample category and at least one sample support image corresponding to the sample query image;
[0143] The training operation includes:
[0144] The sample support image feature corresponding to the sample support image is extracted by the teacher model, the first sample query image feature corresponding to the sample query image is extracted by the teacher model, and the second sample query image feature corresponding to the sample query image is extracted by the initial student model;
[0145] Based on the sample support image feature, a preset number of sample prototype vectors corresponding to the sample support image feature are extracted;
[0146] The second sample query image feature and the preset number of sample prototype vectors are fused to obtain a sample fusion feature, and the total training loss is determined based on the sample fusion feature and the first sample query image feature;
[0147] Based on the total training loss, the parameters of the initial student model are adjusted, and the initial student model after the adjustment is taken as the initial student model corresponding to the next training operation.
[0148] Specifically, the initial student model can be trained based on the trained teacher model and multiple sample sets, wherein each sample set includes a sample query image for a sample category and at least one support image corresponding to the sample query image, the sample category can be a category of a product included in an image, one sample set corresponds to one sample category, the sample query image and the sample support image in one sample set both correspond to the same sample category, and the sample categories corresponding to the multiple sample sets can be different.
[0149] For example, the sample category corresponding to the sample set A is “nut”, and the sample query image and the sample image in the sample set A are both images containing “nut”; the sample category corresponding to the sample set B is “bolt”, and the sample query image and the sample support image in the sample set B are both images containing “bolt”.
[0150] The sample query image and the sample support image in each sample set are both defect-free images of the sample category, that is, the product of the sample category included in the sample query image and the sample support image is defect-free.
[0151] The sample images included in the multiple sample sets can be collected by an image collection device, or can be collected through a network under the premise of complying with relevant regulations, and the present disclosure does not limit this.
[0152] For each sample set, the sample query image and the sample support image in the sample set can be input to the teacher model, and the sample query image and the sample support image are respectively subjected to feature extraction by the teacher model to obtain corresponding first sample query image features and sample support image features.
[0153] In the process of forward knowledge distillation, the sample query image is input to the initial student model, and the sample query image is subjected to multi-scale feature extraction by the plurality of cascaded second network layers in the initial student model to obtain second sample query image features corresponding to each level respectively; in the process of reverse knowledge distillation, the first sample query image features output by the teacher model are input to the initial student model, and the first sample query image features are sequentially subjected to feature reconstruction by the plurality of cascaded second network layers in the initial student model to obtain second sample query image features corresponding to each level respectively.
[0154] Based on the sample support image features, a preset number of sample prototype vectors corresponding to the sample support image features are extracted, and the second sample query image features are fused with the preset number of sample prototype vectors to obtain sample fusion features; and the total training loss is determined based on the sample fusion features and the first sample query image features. The preset number of sample prototype vectors can be used to represent information of a defect-free object of a sample category.
[0155] The parameters of the initial student model are adjusted based on the total training loss, and the initial student model after the adjustment of the parameters is used as the initial student model corresponding to the next training operation. By continuously performing the above training operation, the training of the initial student model is constrained based on the total training loss, so that the image features output by the initial student model and the image features output by the teacher model become more and more close, thereby learning the ability of the teacher model to extract features of a defect-free image, until a training end condition is met, and the initial student model meeting the training end condition is used as a trained student model.
[0156] The training end condition can be that the total training loss converges, for example, the total training loss is less than a set value or the total training loss calculated for a continuous set number of times is less than the set value; or the training end condition can be that the number of training reaches a preset number, which is not limited in the present disclosure.
[0157] In the embodiments of the present disclosure, the teacher model obtains a preset number of sample prototype vectors for representing information of a defect-free object based on a sample support image, and corrects the second sample query image feature extracted by the initial student model towards the direction of the defect-free object by fusing the second sample query image feature with the preset number of sample prototype vectors, thereby effectively suppressing the defect information of the sample query image extracted by the initial student model, so that the feature extracted by the initial student model is as close as possible to the feature of the defect-free image corresponding to the sample query image, and the ability of the trained student model to extract the feature of the defect-free image is improved.
[0158] Further, in the embodiments of the present disclosure, only a small number of defect-free images of different sample categories are needed to effectively train the initial student model, a large number of labeled sample images are not needed, the present disclosure can be applied to different production scenarios, and the manpower and material resources required for labeling samples are saved, and the cost of defect detection is reduced.
[0159] Further, in the embodiments of the present disclosure, the initial student model is trained based on a plurality of sample sets corresponding to different sample categories, so that the initial student model can extract more general image features, and the generalization of the trained student model is improved.
[0160] As an optional embodiment, the total training loss is determined based on the sample fusion feature and the first sample query image feature, including:
[0161] For each level, the second network layer of the level is used to determine the corresponding second sample query image feature based on the sample fusion feature;
[0162] For each level, the first training loss of the level is determined based on the difference between the first sample query image feature corresponding to the first network layer of the level and the second sample query image feature corresponding to the second network layer of the level;
[0163] The total training loss is determined based on the first training loss corresponding to each level.
[0164] Specifically, the teacher model can include a plurality of cascaded first network layers, and the initial student model can include a plurality of cascaded second network layers. It can be understood that the plurality of second network layers in the initial student model can have the same network structure as the plurality of second network layers in the student model, but the parameters of the network have not been trained and learned.
[0165] For each level, the sample support image feature output by the first network layer of the previous level and the first sample query image feature are obtained, and the second sample query image feature output by the second network layer of the previous level is obtained.
[0166] The sample support image feature corresponding to a preset number of sample prototype vectors is extracted, and the second sample query image feature output by the previous level is fused with the preset number of sample prototype vectors to obtain a sample fusion feature.
[0167] The sample fusion feature is input into the second network layer of the level, and the sample fusion feature is further processed by the second network layer to obtain the second sample query image feature output by the second network layer of the level.
[0168] The first sample query image feature output by the first network layer of the previous level is input into the first network layer of the level, and the sample query image feature of the previous level is further extracted by the first network layer to obtain the first sample image query feature output by the first network layer of the level.
[0169] Based on the difference between the first sample query image feature output by the first network layer of the level and the second sample query image feature output by the second network layer of the level, the first training loss of the level is determined.
[0170] Based on the first training loss corresponding to each level, the total training loss of the initial student model is determined. For example, the first training loss corresponding to each level can be directly added to obtain the total training loss, or the first training loss corresponding to each level can be weighted and summed to obtain the total training loss, which is not limited in the embodiments of the present disclosure.
[0171] As an optional embodiment, for each level, based on the difference between the first sample query image feature corresponding to the first network layer of the level and the second sample query image feature corresponding to the second network layer of the level, the first training loss of the level is determined, including:
[0172] Based on the difference between the first sample query image feature corresponding to the first network layer of the level and the second sample query image feature corresponding to the second network layer of the level, the second training loss is determined.
[0173] Based on the orthogonality between the preset number of sample prototype vectors corresponding to the level, the third training loss is determined.
[0174] Based on the second training loss and the third training loss, the first training loss of the level is determined.
[0175] Specifically, for each level, the second training loss can be determined based on the difference between the first sample query image feature output by the first network layer of the level and the second sample query image feature output by the second network layer of the level.
[0176] A third training loss is determined based on the orthogonality between a preset number of sample prototype vectors corresponding to the level, wherein the preset number of sample prototype vectors corresponding to the level are obtained based on sample support image features output by the first network layer of the previous level.
[0177] Optionally, the third training loss may be calculated using an ORTH (Orthogonality Regularization) loss function based on the orthogonality between a preset number of sample prototype vectors corresponding to the level.
[0178] In the embodiment of the present disclosure, a third training loss is obtained by calculating the correlation between a preset number of sample prototype vectors at each level, so as to constrain the orthogonality between the generated preset number of prototype vectors through the third training loss, reduce the linear correlation between the extracted preset number of prototype vectors, and enable the preset number of prototype vectors to more effectively represent the features of the sample support image.
[0179] As an optional embodiment, for each level, the method further includes:
[0180] For each sample support image, determining reference information based on a correlation between a sample support image feature corresponding to the sample support image and a first sample query image feature corresponding to a target sample query image; the target sample query image and the sample support image belong to the same sample set;
[0181] Determining information to be compared based on a correlation between a sample support image feature corresponding to the sample support image and a second sample query image feature corresponding to the sample query image in each sample set;
[0182] Determining a fourth training loss corresponding to the sample support image based on the difference between the reference information and the information to be compared;
[0183] Determine a fifth training loss based on the fourth training losses corresponding to each sample support image;
[0184] The first training loss is determined based on at least one of the following:
[0185] Determine a first training loss based on the second training loss and the fifth training loss;
[0186] A first training loss is determined based on the second training loss, the third training loss, and the fifth training loss.
[0187] Specifically, when a sample query image corresponds to at least two sample support images, for each sample support image, a sample query image that belongs to the same sample set as the sample support image may be used as a target sample query image.
[0188] For each level, the sample support image features of the first network layer output of the level for the sample support image are obtained, and the first sample query image features corresponding to the target sample query image are obtained. The teacher model has better feature expression ability, so the sample support image features and the target sample query image features can be used as a reference sample to represent the correlation of image semantics of different regions between the defect-free images. The reference information is determined based on the correlation between the sample support image features and the first sample query image features corresponding to the target sample query image.
[0189] The second sample query features corresponding to the sample query image in each sample set are obtained based on the second network layer output of the level. The to-be-compared information is determined based on the correlation between the sample support image features corresponding to the sample support image and the second sample query image corresponding to each sample query image in the sample set.
[0190] The fourth training loss corresponding to the sample support image is determined based on the difference between the reference information and the to-be-compared information. The fifth training loss is determined based on the fourth training loss corresponding to each sample support image.
[0191] On this basis, the first training loss of a level can be determined based on the second training loss and the fifth training loss. The first training loss of a level can also be determined based on the second training loss, the third training loss, and the fifth training loss.
[0192] In the embodiments of the present disclosure, for each sample support image, based on the correlation between the sample support image features corresponding to the sample support image and the first sample query image features corresponding to the target sample query image; based on the correlation between the sample support image features corresponding to the sample support image and the second sample query image features corresponding to the sample query image in each sample set, the to-be-compared information is determined; based on the difference between the reference information and the to-be-compared information, the fourth training loss corresponding to the sample support image is determined; based on the fourth training loss corresponding to each sample support image, the fifth training loss is determined, which constrains the difference between the reference information and the to-be-compared information through the fifth training loss, uses the reference information as an anchor point to supervise the initial student model, which helps to optimize the student model; and when each sample query image corresponds to multiple sample support images, the initial student model can be supervised by multiple anchor points in the training process, which is conducive to the robustness of knowledge distillation.
[0193] As an optional embodiment, for each level, the to-be-compared information is determined based on the correlation between the sample support image features corresponding to the sample support image and the second sample query image features corresponding to the sample query image in each sample set, including:
[0194] determine the first to-be-compared information based on a correlation between the sample support image feature corresponding to the sample support image and a second sample query image feature corresponding to the target query image;
[0195] determine the second to-be-compared information based on a correlation between the sample support image feature corresponding to the sample support image and a second sample query image feature corresponding to a non-target sample query image;
[0196] determine the fourth training loss corresponding to the sample support image based on a difference between the reference information and the to-be-compared information, including:
[0197] determine the fourth training loss corresponding to the sample support image based on a difference between the reference information and the first to-be-compared information and a difference between the reference information and the second to-be-compared information.
[0198] Specifically, for the sample query images included in the plurality of sample sets, a sample query image belonging to the same sample set as the sample support image is taken as a target sample query image, and a sample query image belonging to a different sample set as the sample support image is taken as a non-target sample query image, that is, the target sample query image and the sample support image correspond to the same sample class, and the non-target sample query image and the sample support image correspond to different sample classes.
[0199] Therefore, the sample support image feature of the sample support image and the second sample query image feature corresponding to the target sample query image can be taken as a positive sample pair, and the sample support image feature of the sample support image and the second sample query image feature corresponding to the non-target sample query image can be taken as a negative sample pair. The first to-be-compared information is determined based on the correlation between the sample support image feature of the sample support image and the second sample query image feature corresponding to the target sample query image; and the second to-be-compared information is determined based on the correlation between the sample support image feature of the sample support image and the second sample query image feature corresponding to the non-target sample query image.
[0200] The fourth training loss corresponding to a sample support image is determined based on a difference between the reference information and the first to-be-compared information and a difference between the reference information and the second to-be-compared information, so that the initial student model can be constrained to learn that the similarity between the positive sample and the reference sample is increasingly large, and the similarity between the negative sample and the reference sample is increasingly small, thereby improving the robustness of the model.
[0201] Optionally, a specific formula of the fourth training loss is as follows:
[0202]
[0203] wherein, L NCEDenotes the NCE (Noise Contrastive Estimation) loss function, and the sample support image feature for the sample support image at the i-th level is The first query image feature for the sample query image at the i-th level is Will and The correlation matrix between As reference information, the second query image feature for the target query image at the i-th level is The second query image feature for the non-target query image at the i-th level is Will and The correlation matrix between them is used as the first information to be compared Will and The correlation matrix between them is used as the second information to be compared
[0204] The correlation matrix between image features can be calculated based on the following formula:
[0205]
[0206] Where, for and The correlation matrix between them, each position in the correlation matrix represents a feature In (h a ,w a ) and the characteristic In (h b ,w b ) is the cosine similarity of the vectors.
[0207] As an optional embodiment, Figure 7 A flow chart of a model training method provided in an embodiment of the present disclosure is shown as follows: Figure 7 As shown, the model training method provided by the embodiment of the present disclosure is based on the paradigm of reverse knowledge distillation. The teacher model includes multiple cascaded encoders, and the initial student model includes multiple cascaded decoders. For example, the teacher model and the initial student model can be constructed based on the network structure of ResNet (Residual Network).
[0208] During the training process, the sample query image X q Its corresponding sample support image X sThe input is sent to the teacher model, and multi-scale feature extraction is performed on the sample query image and the sample support image through multiple cascaded encoders in the teacher model. Then, the prototype vector is extracted and integrated based on the prototype extraction and integration module (PEIM). Figure 7 The teacher encoder at stage i (i.e., the first network layer at level i) outputs the first sample query image feature corresponding to the sample query image and the sample support image features of the sample support image The characteristic size is C i ×H i ×W i ; Use a convolution (conv) of size 3×3: And a deformation operation (reshape) to get the feature The feature size is L×H i ×E i ;right Use the softmax function to generate L attention maps; each attention map is Multiply element by element and pass the result along Spatial dimension aggregation, thereby generating L prototype vectors That is the process of generating prototype vectors.
[0209] After obtaining L prototype vectors Afterwards, each prototype vector is calculated The second query image features corresponding to the sample query image output by the student decoder at stage i (i.e., the second network layer at level i) Cosine similarity at position (h,w) The feature size is L×H i ×W i ; For position (h,w), select the prototype vector with the greatest similarity as the guiding feature in Will Add the similarity information along the channel direction to obtain the probability map (i.e. the first fusion feature). Finally, the second query image feature Guide Features And the probability plot The splicing will be performed along the channel direction to obtain the splicing result The characteristic size is: (2C i +1)×H i ×W i , which is the process of integrating prototype vectors.
[0210] Will input into the student decoder of the i+1th stage (i.e., the second network layer of the i+1th level), to obtain the corresponding second query image feature of the student decoder output of the i+1th stage input into the teacher encoder of the i+1th stage (i.e., the first network layer of the i+1th level), to obtain the corresponding first query image feature of the teacher encoder output of the i+1th stage input into the teacher encoder of the i+1th stage (i.e., the first network layer of the i+1th level), to obtain the corresponding first query image feature of the teacher encoder output of the i+1th stage based on the difference between and , a training loss of a stage is constructed by using a contrastive distillation strategy (CDS), a total training loss of the model is constructed based on the training loss of each stage respectively, and the training of the model is constrained based on the total training loss.
[0211] wherein the training loss of the i-th stage can be constructed based on the difference between the second query image feature of the student decoder output of the i-th stage and the first query image feature of the teacher encoder output of the i-th stage , and the specific process can be referred to the similar process of the i+1th stage, which will not be described herein.
[0212] wherein the specific formula of the total training loss can be as follows:
[0213]
[0214] wherein i represents the i-th stage, r represents the total number of stages, represents the second training loss constructed based on and represents the fourth training loss constructed based on and represents the support image feature of the n-th sample support image output by the teacher encoder of the i+1th stage, represents the fifth training loss determined based on the k sample support images; represents the third training loss determined based on the orthogonality between the L prototype vectors and
[0215] It should be noted that the above total training loss is explained in units of one sample query image to explain the specific form of the total training loss, and in the actual training process, the final total training loss of the model can also be determined by the total training loss corresponding to the plurality of sample query images in each sample set.
[0216] The model training method provided by the embodiments of the present disclosure proposes a new prototype perception contrast knowledge distillation framework, which utilizes a prototype extraction and integration module (PEIM) to improve the generalization ability of the student model by integrating prior information of a given class from the teacher model into the student model. The PEIM is trained to generate prototypes from a small number of normal samples to provide prior information, and further utilizes these prototypes to guide the student model to reconstruct the distillation target. Subsequently, a contrast distillation strategy is adopted to attract the correlation between samples of the same class while repelling the correlation between samples of different classes, and by constructing a negative sample-positive sample pair as a constraint on the feature correlation between the teacher model and the student model, the normal sample representation and the relationship between samples are distilled, thereby improving the robustness of the model. Only a small number of defect-free images are needed to effectively train the model, thereby overcoming the problems of the prior art that a large number of sample category data are needed for training and the effect is poor under a small number of samples.
[0217] Figure 8 A structural schematic diagram of an image processing apparatus provided by the embodiments of the present disclosure is shown in FIG. 1, which can include: Figure 8
[0218] An image acquisition module 210 is configured to acquire a query image for a target category to be detected and a support image corresponding to the query image;
[0219] A first feature acquisition module 220 is configured to acquire first query image features of the query image corresponding to each level respectively by a plurality of cascaded first network layers of a trained teacher model, and acquire support image features of the support image corresponding to each level respectively;
[0220] A second feature acquisition module 230 is configured to acquire second query image features of the query image corresponding to each level respectively by a plurality of cascaded second network layers of a trained student model; the second image features represent image features obtained after correcting the query image;
[0221] A feature comparison module 240 is configured to determine a first target image feature from the first query image features of the plurality of levels, and determine a second target image feature from the second query image features of the plurality of levels; based on a comparison result between the first target image feature and the second target image feature, a defect detection result for the query image is determined;
[0222] wherein, for each level, the second query image feature is obtained by:
[0223] extracting a preset number of prototype vectors based on the support image features of the previous level; the preset number of prototype vectors are used to represent information of the defect-free objects of the target category;
[0224] fusing the second query image feature of the previous level and the preset number of prototype vectors to obtain a fusion feature, and inputting the fusion feature into the second network layer of the corresponding level to obtain the second query image feature of the level.
[0225] As an optional embodiment, the device further comprises a fusion feature determination module, configured to:
[0226] determine a plurality of similarity information between the second query image feature of the previous level and the preset number of prototype vectors, respectively;
[0227] fuse the plurality of similarity information to obtain a first fusion feature;
[0228] select the prototype vector with the largest similarity to the second query image feature of the previous level as a guide feature;
[0229] fuse the second query image feature of the previous level, the first fusion feature and the guide feature to obtain the fusion feature.
[0230] As an optional embodiment, the device further comprises a model training module, configured to:
[0231] perform at least one training operation on the initial student model based on a plurality of sample sets and the trained teacher model until a training end condition is met, and select the initial student model that meets the training end condition as the trained student model;
[0232] The sample set comprises a sample query image for a sample category and at least one sample support image corresponding to the sample query image;
[0233] wherein, the training operation comprises:
[0234] extracting a sample support image feature corresponding to the sample support image through the teacher model, extracting a first sample query image feature corresponding to the sample query image through the teacher model, and extracting a second sample query image feature corresponding to the sample query image through the initial student model;
[0235] extracting a preset number of sample prototype vectors corresponding to the sample support image feature based on the sample support image feature;
[0236] fusing the second sample query image feature with the preset number of sample prototype vectors to obtain a sample fusion feature; determining a total training loss based on the sample fusion feature and the first sample query image feature;
[0237] adjusting parameters of the initial student model based on the total training loss, and taking the initial student model after the adjustment of the parameters as an initial student model corresponding to a next training operation.
[0238] As an optional embodiment, when the model training module determines the total training loss based on the sample fusion feature and the first sample query image feature, the model training module is specifically configured to:
[0239] for each level, determining a corresponding second sample query image feature based on the sample fusion feature through the second network layer of the level;
[0240] for each level, determining a first training loss of the level based on a difference between the first sample query image feature corresponding to the first network layer of the level and the second sample query image feature corresponding to the second network layer of the level;
[0241] determining the total training loss based on the first training loss corresponding to each of the multiple levels.
[0242] As an optional embodiment, for each level, when the model training module determines the first training loss of the level based on a difference between the first sample query image feature corresponding to the first network layer of the level and the second sample query image feature corresponding to the second network layer of the level, the model training module is specifically configured to:
[0243] determining a second training loss based on a difference between the first sample query image feature corresponding to the first network layer of the level and the second sample query image feature corresponding to the second network layer of the level;
[0244] determining a third training loss based on orthogonality between the preset number of sample prototype vectors corresponding to the level;
[0245] determining the first training loss of the level based on the second training loss and the third training loss.
[0246] As an optional embodiment, the apparatus further comprises a contrast learning module configured to:
[0247] for each sample support image, determining reference information based on a correlation between a sample support image feature corresponding to the sample support image and a first sample query image feature corresponding to a target sample query image; the target sample query image and the sample support image belong to a same sample set;
[0248] determine the to-be-compared information based on a correlation between the sample support image features corresponding to the sample support images and second sample query image features corresponding to the sample query images in each sample set;
[0249] determine the fourth training loss corresponding to the sample support images based on a difference between the reference information and the to-be-compared information;
[0250] determine a fifth training loss based on the fourth training loss corresponding to each sample support image respectively;
[0251] The first training loss is determined based on at least one of the following manners:
[0252] determine the first training loss based on the second training loss and the fifth training loss;
[0253] determine the first training loss based on the second training loss, the third training loss and the fifth training loss.
[0254] As an optional embodiment, when determining the to-be-compared information based on a correlation between the sample support image features corresponding to the sample support images and second sample query image features corresponding to the sample query images in each sample set, the contrast learning module is specifically configured to:
[0255] determine first to-be-compared information based on a correlation between the sample support image features corresponding to the sample support images and second sample query image features corresponding to the target query image;
[0256] determine second to-be-compared information based on a correlation between the sample support image features corresponding to the sample support images and second sample query image features corresponding to the non-target sample query image;
[0257] When determining the fourth training loss corresponding to the sample support images based on a difference between the reference information and the to-be-compared information, the contrast learning module is specifically configured to:
[0258] determine the fourth training loss corresponding to the sample support images based on a difference between the reference information and the first to-be-compared information and a difference between the reference information and the second to-be-compared information.
[0259] The apparatuses provided in the embodiments of the present disclosure can perform the methods provided in the embodiments of the present disclosure, and the implementation principles are similar, and have corresponding technical effects. The actions performed by each module in the apparatuses of the embodiments of the present disclosure are corresponding to the steps in the methods of the embodiments of the present disclosure. For the detailed function description of each module of the apparatus, please refer to the description of the corresponding method in the foregoing description, which will not be repeated here.
[0260] In the embodiments of the present disclosure, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined target, and can be implemented entirely or partially by using software, hardware (such as a processing circuit or a memory) or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an integral module or unit that includes the functions of the module or unit.
[0261] The embodiments of the present disclosure provide an electronic device, including a memory, a processor and a computer program stored in the memory, and the processor executes the computer program to implement the steps of the method provided by any of the optional embodiments of the present disclosure. Compared with the prior art, the following can be achieved: a preset number of prototype vectors for representing information representing a defect-free object are obtained based on a support image by a teacher model, the second query image features are fused with the preset number of prototype vectors, and the prototype vectors are used to correct the second query image features extracted by the student model towards the direction of the defect-free object, effectively suppressing the defect information of the query image extracted by the student model, so that the features extracted by the student model are as close as possible to the features of the defect-free image corresponding to the query image, expanding the difference between the query image features output by the student model and the teacher model, effectively detecting defects, and improving the accuracy of the defect detection result.
[0262] In one optional embodiment, an electronic device is provided, as shown in Figure 9 The electronic device 4000 shown in Figure 9 The electronic device 4000 shown in
[0263] The processor 4001 can be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) or other programmable logic device, transistor logic device, hardware component, or any combination thereof. It can implement or execute the various exemplary logical blocks, modules and circuits described in connection with the disclosure. The processor 4001 can also be a combination of computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.
[0264] The bus 4002 can include a path for transmitting information between the above-mentioned components. The bus 4002 can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, or the like. The bus 4002 can be divided into an address bus, a data bus, a control bus, and the like. For convenience of representation, Figure 9 In the figure, only one thick line is used to represent the bus, but it does not mean that there is only one bus or only one type of bus.
[0265] The memory 4003 can be a ROM (Read Only Memory) or other type of static storage device that can store static information and instructions, a RAM (Random Access Memory) or other type of dynamic storage device that can store information and instructions, an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, an optical disk storage (including a compact disk, a laser disk, an optical disk, a digital versatile disk, a Blu-ray disk, and the like), a magnetic disk storage medium, other magnetic storage device, or any other medium capable of carrying or storing computer programs and capable of being read by a computer, without limitation.
[0266] The memory 4003 is configured to store a computer program for implementing the embodiments of the present disclosure, and the processor 4001 is configured to control the execution of the computer program stored in the memory 4003. The processor 4001 is configured to execute the computer program stored in the memory 4003 to implement the steps shown in the foregoing method embodiments.
[0267] The electronic device includes, but is not limited to, a mobile terminal such as a mobile phone, a notebook computer, a digital broadcast receiver, a PDA (Personal Digital Assistant), a PAD (Tablet Personal Computer), a PMP (Portable Multimedia Player), a car terminal (for example, a car navigation terminal), a wearable device, and the like, and a fixed terminal such as a digital TV, a desktop computer, and the like.
[0268] The embodiments of the present disclosure provide a computer readable storage medium, and the computer readable storage medium stores a computer program. When the computer program is executed by a processor, the steps and corresponding contents of the foregoing method embodiments can be implemented.
[0269] The embodiments of the present disclosure also provide a computer program product, and the computer program product includes a computer program. When the computer program is executed by a processor, the steps and corresponding contents of the foregoing method embodiments can be implemented.
[0270] It should be understood that, although the flowcharts of the embodiments of the present disclosure indicate the respective operation steps by arrows, the implementation order of the steps is not limited to the order indicated by the arrows. Unless otherwise specified herein, in some implementation scenarios of the embodiments of the present disclosure, the implementation steps in each flowchart can be executed in other orders as required. In addition, part or all of the steps in each flowchart can include multiple sub-steps or multiple stages based on the actual implementation scenario. Part or all of the sub-steps or stages can be executed at the same time, and each of the sub-steps or stages can also be executed at different times. In the scenario where the execution times are different, the execution order of the sub-steps or stages can be flexibly configured as required, and the embodiments of the present disclosure do not limit this.
[0271] The above is only an optional implementation of some scenarios of the present disclosure, and it should be pointed out that, for those skilled in the art, other similar implementation manners based on the technical concept of the present disclosure can also be adopted without departing from the technical concept of the present disclosure, and such implementation manners also belong to the protection scope of the embodiments of the present disclosure.
Claims
1. An image processing method, characterized by, The method comprises the following steps: obtaining a query image of a target category to be detected and a support image corresponding to the query image; obtaining first query image features corresponding to the query image and support image features corresponding to the support image through a plurality of cascaded first network layers of a trained teacher model; obtaining second query image features corresponding to the query image through a plurality of cascaded second network layers of a trained student model; the second image features represent image features obtained after the query image is corrected; determining first target image features from the first query image features of the plurality of levels and determining second target image features from the second query image features of the plurality of levels; based on a comparison result between the first target image features and the second target image features, determining a defect detection result for the query image; wherein, for each level, the second query image features are obtained in the following manner: based on the support image features of the previous level, extracting a preset number of prototype vectors; the preset number of prototype vectors are used to represent information of a defect-free object of the target category; fusing the second query image features of the previous level with the preset number of prototype vectors to obtain fusion features, and inputting the fusion features into the second network layer of the corresponding level to obtain the second query image features of the level.
2. The image processing method of claim 1, wherein, The fusion of the second query image features of the previous level with the preset number of prototype vectors to obtain the fusion features comprises: determining a plurality of similarity information between the second query image features of the previous level and the preset number of prototype vectors; fusing the plurality of similarity information to obtain first fusion features; taking the prototype vector with the maximum similarity with the second query image features of the previous level as a guide feature; fusing the second query image features of the previous level, the first fusion features, and the guide feature to obtain the fusion features.
3. The image processing method of claim 1, wherein, The student model is trained in the following manner: based on a plurality of sample sets and a trained teacher model, performing at least one training operation on an initial student model until a training end condition is met, and taking the initial student model that meets the training end condition as a trained student model; the sample set comprises sample query images of a sample category and at least one sample support image corresponding to the sample query images; wherein, the training operation comprises: extracting sample support image features corresponding to the sample support image through the teacher model, extracting first sample query image features corresponding to the sample query image through the teacher model, and extracting second sample query image features corresponding to the sample query image through the initial student model; based on the sample support image features, extracting a preset number of sample prototype vectors corresponding to the sample support image features; fusing the second sample query image features with the preset number of sample prototype vectors to obtain sample fusion features; based on the sample fusion features and the first sample query image features, determining a total training loss; Adjust parameters of the initial student model based on the total training loss, and take the initial student model after adjusting the parameters as an initial student model corresponding to a next training operation.
4. The image processing method of claim 3, wherein, The total training loss is determined based on the sample fusion feature and the first sample query image feature. For each level, a second sample query image feature corresponding to the level is determined based on the sample fusion feature and the second network layer of the level. For each level, a first training loss of the level is determined based on a difference between the first sample query image feature corresponding to the first network layer of the level and the second sample query image feature corresponding to the second network layer of the level. The total training loss is determined based on the first training loss corresponding to each of the plurality of levels.
5. The image processing method of claim 4, wherein, For each level, the first training loss of the level is determined based on a difference between the first sample query image feature corresponding to the first network layer of the level and the second sample query image feature corresponding to the second network layer of the level. A second training loss is determined based on the difference between the first sample query image feature corresponding to the first network layer of the level and the second sample query image feature corresponding to the second network layer of the level. A third training loss is determined based on orthogonality between the preset number of sample prototype vectors corresponding to the level. The first training loss of the level is determined based on the second training loss and the third training loss.
6. The image processing method of claim 5, further comprising: For each sample support image, reference information is determined based on a correlation between a sample support image feature corresponding to the sample support image and the first sample query image feature corresponding to the target sample query image. The target sample query image and the sample support image belong to a same sample set. To-be-compared information is determined based on a correlation between the sample support image feature corresponding to the sample support image and a second sample query image feature corresponding to a sample query image in each sample set. A fourth training loss corresponding to the sample support image is determined based on a difference between the reference information and the to-be-compared information. A fifth training loss is determined based on the fourth training loss corresponding to each sample support image. The first training loss is determined based on at least one of the following: The first training loss is determined based on the second training loss and the fifth training loss. The first training loss is determined based on the second training loss, the third training loss, and the fifth training loss.
7. The image processing method of claim 6, wherein the to-be-compared information is determined based on a correlation between the sample support image feature corresponding to the sample support image and a second sample query image feature corresponding to a target query image. First to-be-compared information is determined based on a correlation between the sample support image feature corresponding to the sample support image and the second sample query image feature corresponding to the target query image. determine, based on a correlation between the sample support image feature corresponding to the sample support image and a second sample query image feature corresponding to the non-target sample query image, second to-be-compared information; The fourth training loss corresponding to the sample support image is determined based on the difference between the reference information and the to-be-compared information, including: The fourth training loss corresponding to the sample support image is determined based on the difference between the reference information and the first to-be-compared information, and the difference between the reference information and the second to-be-compared information.
8. An image processing apparatus characterized by comprising: including: An image acquisition module is configured to acquire a query image for a target category to be detected and a support image corresponding to the query image; A first feature acquisition module is configured to acquire first query image features corresponding to the query image at respective levels by a plurality of cascaded first network layers of a trained teacher model, and acquire support image features corresponding to the support image at respective levels; A second feature acquisition module is configured to acquire second query image features corresponding to the query image at respective levels by a plurality of cascaded second network layers of a trained student model; the second image features represent image features obtained after the query image is corrected; A feature comparison module is configured to determine a first target image feature from the first query image features at multiple levels, and determine a second target image feature from the second query image features at multiple levels; based on a comparison result between the first target image feature and the second target image feature, a defect detection result for the query image is determined; For each level, the second query image feature is acquired by the following way: Based on the support image feature of the previous level, a preset number of prototype vectors are extracted; the preset number of prototype vectors are used to represent information of a defect-free object of the target category; The second query image feature of the previous level and the preset number of prototype vectors are fused to obtain a fusion feature, and the fusion feature is input into the second network layer of the corresponding level to obtain the second query image feature of the level.
9. An electronic device comprising a memory, a processor, and a computer program stored on the memory, wherein the computer program, when executed by the processor, is arranged to perform the method of any one of claims 1 to 8. The processor executes the computer program to implement the steps of the method of any one of claims 1 to 7.
10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 7.
11. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 7. The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 7.