Image recognition method and apparatus, storage medium and electronic device

By using an abnormal part recognition network with a twin network structure, the feature similarity of industrial part images is extracted and compared, the problem of insufficient recognition accuracy of industrial part images in the prior art is solved, and higher recognition accuracy and detection ability of minor defects are achieved.

WO2025060766A9PCT designated stage expired Publication Date: 2025-05-08TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/112262
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-09-22
Filing Date
2024-08-15
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

The existing industrial parts image recognition methods have shortcomings in the accuracy of the recognition results, especially when the defects on the parts to be tested are small or minor, which can easily lead to missed inspection.

Method used

An abnormal part recognition network with a twin network structure is used, and the image features of the image to be analyzed are extracted through the first feature extraction network and the second feature extraction network, and the feature similarity between the two is calculated to determine whether the part to be detected is an abnormal part.

Benefits of technology

It improves the accuracy of the recognition results of industrial parts image recognition, reduces missed inspection of NG products, and enhances the detection ability of minor defects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024112262_08052025_PF_FP_ABST
    Figure CN2024112262_08052025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present application are an image recognition method and apparatus, a storage medium and an electronic device, which can be applied to various scenarios such as cloud technology, artificial intelligence, smart transportation and driver assistance. The method comprises: using an anomaly part recognition network having a siamese network structure to extract image features of an image to be analyzed, the image features comprising a first image feature extracted by a first feature extraction network in the siamese network structure from said image, and a second image feature extracted by a second feature extraction network in the siamese network structure from said image, the first feature extraction network being trained by means of positive sample images, and the second feature extraction network being trained by means of negative sample images with the first feature extraction network as a reference network; acquiring the feature similarity between the first image feature and the second image feature; and, on the basis of the feature similarity, determining a recognition result of said image. The present application solves the technical problem of relatively low accuracy of recognition results of existing image recognition modes for industrial parts.
Need to check novelty before this filing date? Find Prior Art

Description

Image recognition method and device, storage medium and electronic device

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on September 22, 2023, application number 202311239167.9, and application name “Image Recognition Method and Device, Storage Medium and Electronic Device”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present invention relates to the field of computers, and more specifically, to image recognition technology. Background Art

[0003] In many industrial scenarios, the production and assembly of a single product often requires a large number of industrial parts. During the mass production of these parts, due to their small size, many defective, substandard parts (also known as NG (No Good)) inevitably appear. Therefore, before these parts are used, they must be inspected for abnormal defects.

[0004] Currently, the commonly used abnormality detection method for industrial parts is: collecting images of the part to be inspected and images of normal parts (also called OK parts), and then performing feature comparison based on the two images. If the similarity between the two features is high, the part to be inspected is identified as a OK part; if the similarity between the two features is low, the part to be inspected is identified as an NG part.

[0005] However, when defects on the parts to be inspected are small or minor, their characteristics may not differ significantly from those of approved parts, resulting in missed inspections of defective parts. In other words, the image recognition methods for industrial parts provided in related technologies suffer from low recognition accuracy.

[0006] To address the above-mentioned problems, no effective solutions have been proposed so far.

[0007] Summary of the Invention

[0008] The embodiments of the present application provide an image recognition method and device, a storage medium, and an electronic device to at least solve the technical problem of low accuracy of recognition results in existing image recognition methods for industrial parts.

[0009] According to one aspect of an embodiment of the present application, an image recognition method is provided, which is executed by an electronic device and includes: obtaining an image to be analyzed obtained by image acquisition of a part object to be detected; extracting image features of the image to be analyzed using an abnormal part recognition network having a twin network structure, wherein the image features include first image features extracted from the image to be analyzed by the first feature extraction network in the twin network structure, and second image features extracted from the image to be analyzed by the second feature extraction network in the twin network structure, the first feature extraction network is trained using positive sample images, and the second feature extraction network is trained using negative sample images with the first feature extraction network as a reference network, the positive sample image is a sample image including a normal part object in the sample image pair, and the negative sample image is a sample image including an abnormal part object synthesized based on the positive sample image; obtaining feature similarity between the first image feature and the second image feature; determining a recognition result of the image to be analyzed based on the feature similarity, wherein the recognition result is used to indicate whether the part object to be detected in the image to be analyzed is the abnormal part object.

[0010] According to another aspect of an embodiment of the present application, an image recognition device is also provided, which is deployed on an electronic device and includes: a first acquisition unit, used to acquire an image to be analyzed obtained by image acquisition of a part object to be detected; an identification unit, used to extract image features of the above-mentioned image to be analyzed using an abnormal part recognition network with a twin network structure, wherein the above-mentioned image features include the first image features extracted by the first feature extraction network in the above-mentioned twin network structure for the above-mentioned image to be analyzed, and the second image features extracted by the second feature extraction network in the above-mentioned twin network structure for the above-mentioned image to be analyzed, the above-mentioned first feature extraction network is trained using positive sample images, and the above-mentioned second feature extraction network is trained using negative sample images with the above-mentioned first feature extraction network as a reference network, the positive sample image is a sample image including a normal part object in the sample image pair, and the negative sample image is a sample image including an abnormal part object synthesized based on the positive sample image; a second acquisition unit, used to obtain feature similarity between the above-mentioned first image feature and the above-mentioned second image feature; a determination unit, used to determine the recognition result of the above-mentioned image to be analyzed based on the above-mentioned feature similarity, wherein the above-mentioned recognition result is used to indicate whether the above-mentioned part object to be detected in the above-mentioned image to be analyzed is the above-mentioned abnormal part object.

[0011] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided, in which a computer program is stored. The computer program is configured to execute the above-mentioned image recognition method when running.

[0012] According to another aspect of the embodiments of the present application, a computer program product is provided, the computer program product including a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the above image recognition method.

[0013] According to another aspect of the embodiments of the present application, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the image recognition method through the computer program.

[0014] In an embodiment of the present application, an image to be analyzed is obtained by capturing an image of a part to be inspected. Then, an abnormal part recognition network having a Siamese network structure is used to extract image features from the image to be analyzed. The image features include first image features extracted by a first feature extraction network within the Siamese network structure and second image features extracted by a second feature extraction network within the Siamese network structure. The first feature extraction network is trained using positive sample images, while the second feature extraction network is trained using negative sample images, using the first feature extraction network as a reference network. Positive sample images are sample images from a sample image pair that include a normal part object, while negative sample images are sample images synthesized based on the positive sample images and include an abnormal part object. During the training process, the output of the second feature extraction network is aligned as closely as possible with the output of the first feature extraction network. This ensures that the second feature extraction network is robust to abnormal defects, while the first feature extraction network, which serves as the reference network, is not robust to abnormal defects. Therefore, during recognition, whether the part to be inspected is an abnormal part object can be determined by determining whether the features output by the first and second feature extraction networks are similar. Specifically, the feature similarity between the first and second image features is determined. Furthermore, the recognition result of the image to be analyzed is determined based on the feature similarity, wherein the recognition result is used to indicate whether the part object to be detected in the image to be analyzed is an abnormal part object. That is to say, in the embodiment of the present application, the similarity between the features of the image to be analyzed corresponding to the part object to be detected extracted by each branch network (i.e., the first feature extraction network, the second feature extraction network) in the abnormal part recognition network is used to determine whether the part object to be detected has abnormal defects. It is not a single determination of the recognition result of the image to be analyzed based on the similarity between the image features of the image corresponding to the part object to be detected and the image features of the image corresponding to the normal part object. This solves the technical problem of low accuracy of the recognition results of the image recognition method for industrial parts in the prior art, and achieves the technical effect of improving the accuracy of the recognition results of the image recognition method for industrial parts. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0016] FIG1 is a schematic diagram of an application environment of an optional image recognition method according to an embodiment of the present application;

[0017] FIG2 is a flow chart of an optional image recognition method according to an embodiment of the present application;

[0018] FIG3 is a schematic diagram of an optional image recognition method according to an embodiment of the present application;

[0019] FIG4 is a schematic diagram of an optional image recognition method according to an embodiment of the present application;

[0020] FIG5 is a schematic diagram of an optional image recognition method according to an embodiment of the present application;

[0021] FIG6 is a schematic diagram of an optional image recognition method according to an embodiment of the present application;

[0022] FIG7 is a flowchart of an optional image recognition method according to an embodiment of the present application;

[0023] FIG8 is a schematic diagram of an optional image recognition method according to an embodiment of the present application;

[0024] FIG9 is a schematic diagram of an optional image recognition method according to an embodiment of the present application;

[0025] FIG10 is a flowchart of an optional image recognition method according to an embodiment of the present application;

[0026] FIG11 is a schematic diagram of an optional image recognition method according to an embodiment of the present application;

[0027] FIG12 is a flowchart of an optional image recognition method according to an embodiment of the present application;

[0028] FIG13 is a flowchart of an optional image recognition method according to an embodiment of the present application;

[0029] FIG14 is a schematic structural diagram of an optional image recognition device according to an embodiment of the present application;

[0030] FIG15 is a schematic structural diagram of an optional electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0031] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0032] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar parts and objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0033] According to one aspect of an embodiment of the present application, an image recognition method is provided. Optionally, as an optional implementation, the above-mentioned image recognition method can be, but is not limited to, applied in the environment shown in Figure 1. As shown in Figure 1, the terminal device 102 includes a memory 104 for storing various data generated during the operation of the terminal device 102, a processor 106 for processing and calculating the above-mentioned data, and a display 108 for displaying the image to be analyzed. The terminal device 102 can exchange data with a server 112 through a network 110. The server 112 is connected to a database 114, which is used to store various data. The terminal device 102 can run a program application for identifying the image to be analyzed.

[0034] The specific application process of the above method in the environment shown in Figure 1 is as follows:

[0035] During steps S102 - S104 , the terminal device 102 acquires the image to be analyzed obtained by performing image acquisition on the part to be inspected, and sends the image to be analyzed to the server 112 via the network 110 .

[0036] Then, S106-S110 are executed. Upon receiving the image to be analyzed, server 112 uses an abnormal part recognition network with a Siamese network structure to extract image features of the image to be analyzed. The image features include first image features extracted from the image to be analyzed by a first feature extraction network within the Siamese network structure and second image features extracted from the image to be analyzed by a second feature extraction network within the Siamese network structure. The first feature extraction network is trained using positive sample images, and the second feature extraction network is trained using negative sample images, using the first feature extraction network as a reference network. The positive sample image is a sample image in a sample image pair that includes a normal part object, while the negative sample image is a sample image synthesized based on the positive sample image that includes an abnormal part object. Server 112 obtains feature similarity between the first image feature and the second image feature. Based on the feature similarity, server 112 determines a recognition result for the image to be analyzed, wherein the recognition result indicates whether the part object to be detected in the image to be analyzed is an abnormal part object.

[0037] Then, S112 is executed, and the server 112 network 110 sends the recognition result of the image to be analyzed to the terminal device 102.

[0038] In an embodiment of the present application, an image to be analyzed is obtained by capturing an image of a part object to be detected. Then, an abnormal part object recognition network having a twin network structure is used to extract image features of the image to be analyzed, wherein the image features include first image features extracted from the image to be analyzed by a first feature extraction network in the twin network structure, and second image features extracted from the image to be analyzed by a second feature extraction network in the twin network structure. The first feature extraction network is trained using positive sample images, and the second feature extraction network is trained using negative sample images with the first feature extraction network as a reference network. The positive sample image is a sample image in a sample image pair that includes a normal part object, and the negative sample image is a sample image synthesized based on the positive sample image that includes an abnormal part object. Feature similarity between the first image feature and the second image feature is obtained. Furthermore, based on the feature similarity, a recognition result for the image to be analyzed is determined, wherein the recognition result indicates whether the part object to be detected in the image to be analyzed is an abnormal part object. That is to say, in the embodiment of the present application, the similarity between the features of the image to be analyzed corresponding to the part to be detected, which are extracted by each branch network (i.e., the first feature extraction network and the second feature extraction network) in the abnormal part object recognition network, is used to determine whether the part to be detected has an abnormal defect. Instead of simply determining the recognition result of the image to be analyzed based on the similarity between the image features of the image corresponding to the part to be detected and the image features of the image corresponding to the normal part object, this solves the technical problem of low recognition accuracy of the image recognition method for industrial parts in the prior art, and achieves the technical effect of improving the recognition accuracy of the image recognition method for industrial parts.

[0039] In this embodiment, the above-mentioned terminal device can be a terminal device configured with a client, which can include but is not limited to at least one of the following: a mobile phone (such as an Android phone, an iOS phone, etc.), a laptop computer, a tablet computer, a PDA, an MID (Mobile Internet Devices), a PAD, a desktop computer, an intelligent voice interaction device, a smart home appliance, a car terminal, etc. The client can be a video client, an instant messaging client, a browser client, an education client, etc. The above-mentioned network can include but is not limited to: a wired network, a wireless network, wherein the wired network includes: a local area network, a metropolitan area network and a wide area network, and the wireless network includes: Bluetooth, WIFI and other networks that realize wireless communication. The above-mentioned server can be a single server, or it can be a server cluster composed of multiple servers, or a cloud server. The above is only an example, and this embodiment does not impose any limitation on this.

[0040] As an optional solution, as shown in FIG2 , the image recognition method includes:

[0041] S202, acquiring an image to be analyzed obtained by collecting an image of the part to be inspected;

[0042] It is understood that the above-described image recognition method can be applied, but is not limited to, to industrial parts anomaly detection scenarios. Specifically, the part to be detected can be a part object requiring determination of abnormal defects, or any other part requiring inspection. When the image recognition method is applied to industrial parts anomaly detection, the part to be detected can be an industrial part; when the image recognition method is applied to automotive parts anomaly detection, the part to be detected can be an automotive part; and when the image recognition method is applied to mechanical parts anomaly detection, the part to be detected can be a mechanical part. The image to be analyzed can be an image acquired by capturing the part to be detected, so that image features included in the image can be identified through subsequent analysis. A first feature extraction network in the abnormal part recognition network extracts first image features of the image to be analyzed corresponding to the part to be detected, and a second feature extraction network in the abnormal part recognition network extracts second image features of the image to be analyzed corresponding to the part to be detected. Based on the feature similarity between the first and second image features, it is determined whether the part to be detected is an abnormal part. In addition, the above-mentioned image recognition method can also be applied to the classification of any other part objects besides industrial parts, or in abnormality detection scenarios, which is not limited in this embodiment.

[0043] Assuming the above-mentioned image recognition method is applied to an industrial parts anomaly detection scenario, the above-mentioned part to be detected can be, but is not limited to, used to indicate the industrial part to be detected, such as a screw, gear, chain, pulley, etc., and this application does not impose any restrictions on this. The image to be analyzed obtained by capturing the image of the above-mentioned part to be detected can be, but is not limited to, images obtained by capturing the image of the industrial part to be detected from various angles.

[0044] In this embodiment, after acquiring the image to be analyzed by capturing the image of the part to be detected, the method may further include, but is not limited to, correcting the display position of the part to be detected in the image to be analyzed using a template image corresponding to the part to be detected, thereby obtaining a corrected image to be analyzed. The template image may include a reference part object that belongs to the same category as the part to be detected.

[0045] S204, extracting image features of the image to be analyzed using an abnormal part recognition network having a twin network structure, wherein the image features include first image features extracted from the image to be analyzed by a first feature extraction network in the twin network structure, and second image features extracted from the image to be analyzed by a second feature extraction network in the twin network structure, the first feature extraction network being trained using positive sample images, and the second feature extraction network being trained using negative sample images with the first feature extraction network as a reference network, the positive sample image being a sample image in a sample image pair including a normal part object, and the negative sample image being a sample image synthesized based on the positive sample image including an abnormal part object;

[0046] A twin network structure is a type of neural network structure that consists of two or more identical networks. Twin network structures are often used to solve similarity-based tasks. In the embodiments of this application, twin network structures mainly refer to paired structures.

[0047] It should be noted that the abnormal part recognition network with a twin network structure can be used, but is not limited to, to indicate a distilled neural network consisting of a teacher network and a student network. The initial student network has the same network structure as the teacher network. Assuming that the abnormal part recognition network with a twin network structure is a distilled neural network consisting of a teacher network and a student network, the first feature extraction network can be used, but is not limited to, to indicate the teacher network, and the second feature extraction network can be used, but is not limited to, to indicate the student network. During the training process of the abnormal part recognition network (i.e., the distilled neural network), the first feature extraction network (i.e., the teacher network) is used to extract features from positive sample images showing normal part objects to obtain positive sample image features. Then, the second feature extraction network (i.e., the student network) is trained using negative sample images, and the parameters of the second feature extraction network (i.e., the student network) are continuously adjusted until the loss function between the negative sample image features extracted by the second feature extraction network (i.e., the student network) and the positive sample image features reaches a predetermined threshold. When the loss function between the negative sample image features and the positive sample image features reaches a predetermined threshold, the abnormal part recognition network (i.e., the distilled neural network) is considered trained. The goal of training is to make the output of the second feature extraction network as close as possible to the output of the first feature extraction network. This makes the second feature extraction network robust to abnormal defects, while the first feature extraction network, which serves as the reference network, is not robust to abnormal defects.

[0048] Furthermore, when using the abnormal part recognition network to analyze an image to be analyzed, the image to be analyzed can be input into the abnormal part recognition network (i.e., a distilled neural network). If the image to be analyzed is a second-category image (i.e., a normal image), the image features output by the teacher network and the student network are relatively similar, or even identical. If the image to be analyzed is a first-category image (i.e., an abnormal image), the image features output by the teacher network are similar to those of the negative sample image, while the image features output by the student network, due to its robustness, are more similar to those of the positive sample image. Therefore, the image features output by the two networks differ significantly. In other words, if the image features extracted by the first feature extraction network (i.e., the teacher network) in the abnormal part recognition network (i.e., the distilled neural network) are the same as the image features extracted by the second feature extraction network (i.e., the student network) in the abnormal part recognition network (i.e., the distilled neural network), then the part object included in the image to be analyzed and the part object included in the positive sample image belong to the same category. Accordingly, if the image features extracted by the first feature extraction network (i.e., the teacher network) are different from the image features extracted by the second feature extraction network (i.e., the student network), it means that the part object included in the above-mentioned image to be detected and the part object included in the negative sample image belong to the same category.

[0049] In this embodiment, the first feature extraction network can also be trained using positive sample images of normal part objects from the sample image pairs, and the second feature extraction network can also be trained using positive and negative sample images from the sample image pairs. For example, assuming the abnormal part recognition network with a twin network structure is a distillation neural network consisting of a teacher network and a student network, the first feature extraction network is the teacher network, and the second feature extraction network is the student network. The training process for the abnormal part recognition network (i.e., the distillation neural network) includes the following steps: inputting a positive sample image of a normal part object into the first feature extraction network (i.e., the teacher network), obtaining image features extracted from the positive sample image by the first feature extraction network (i.e., the teacher network), and comparing these image features with the actual image features of the positive sample image. The above steps are repeated, continuously adjusting the network parameters of the first feature extraction network (i.e., the teacher network) until the loss function between the image features extracted from the positive sample image by the first feature extraction network (i.e., the teacher network) and the actual image features of the positive sample image is less than a predetermined threshold. For the second feature extraction network (i.e., the student network), a positive sample image of a normal part object is input into the second feature extraction network (i.e., the student network). The image features extracted by the second feature extraction network (i.e., the student network) for the positive sample image are obtained. These image features are then compared with the image features extracted by the trained first feature extraction network (i.e., the teacher network) for the positive sample image. Furthermore, a negative sample image of an abnormal part object is input into the second feature extraction network (i.e., the student network). The image features extracted by the second feature extraction network (i.e., the student network) for the negative sample image are obtained. These image features are then compared with the image features extracted by the trained first feature extraction network (i.e., the teacher network) for the negative sample image. The above steps are repeated continuously to continuously adjust the network parameters of the second feature extraction network (i.e., the student network). Until the loss function between the image features extracted by the second feature extraction network (i.e., the student network) for the positive sample image and the image features extracted by the first feature extraction network (i.e., the teacher network) for the positive sample image, and the image features extracted by the second feature extraction network (i.e., the student network) for the negative sample image and the image features extracted by the first feature extraction network (i.e., the teacher network) for the negative sample image is less than a predetermined threshold.

[0050] Assume that the first feature extraction network is trained using positive sample images including normal part objects in the sample image pairs, and the second feature extraction network is trained using positive sample images and negative sample images in the sample image pairs. In this embodiment, before acquiring the image to be analyzed obtained by image acquisition of the part object to be detected, the method further includes: S1, acquiring K positive sample images including normal part objects and L negative sample images including abnormal part objects, where K is a natural number greater than 1 and L is a natural number greater than 1; S2, training the initialized abnormal part object recognition network using the K positive sample images and the L negative sample images until the comparison result between the output features of the second feature extraction network and the output features of the first feature extraction network reaches a convergence condition.

[0051] Optionally, the aforementioned method of training the initialized abnormal part recognition network using K positive sample images and L negative sample images until the comparison result between the output features of the second feature extraction network and the output features of the first feature extraction network reaches a convergence condition may further include: training the initialized abnormal part recognition network using K positive sample images and L negative sample images, and if the comparison result between the output features of the second feature extraction network and the output features of the first feature extraction network does not reach a convergence condition, adjusting relevant parameters in the training of the abnormal part recognition network. Then, training the adjusted abnormal part recognition network continues until the comparison result between the output features of the second feature extraction network and the output features of the first feature extraction network reaches a convergence condition.

[0052] Optionally, the above-mentioned method of using K positive sample images and L negative sample images to train the initialized abnormal part recognition network until the comparison result between the output features of the second feature extraction network and the output features of the first feature extraction network reaches a convergence condition can also include: using K positive sample images to train the first feature extraction network in the initialized abnormal part recognition network until a first sub-convergence condition is reached; inputting the K positive sample images and L negative sample images into the first feature extraction network that reaches the first sub-convergence condition and the second feature extraction network in the initialized abnormal part recognition network for training until a second sub-convergence condition is reached. Among them, after the positive sample image is input into the first feature extraction network that meets the first sub-convergence condition, the first sample sub-image feature is obtained, and after the positive sample image is input into the second feature extraction network, the second sample sub-image feature is obtained. After the negative sample image is input into the first feature extraction network that meets the first sub-convergence condition, the third sample sub-image feature is obtained, and after the negative sample image is input into the second feature extraction network, the fourth sample sub-image feature is obtained. The second sub-convergence condition indicates that the feature similarity between the first sample sub-image feature and the second sample sub-image feature is greater than the seventh threshold, but the feature similarity between the third sample sub-image feature and the fourth sample sub-image feature is less than the eighth threshold.

[0053] It should be noted that the above-mentioned acquisition of L negative sample images showing abnormal part objects may include, but is not limited to: acquiring K positive sample images showing normal part objects; and adjusting the L positive sample images obtained from the K positive sample images to obtain L negative sample images.

[0054] Optionally, the above-mentioned training of the initialized abnormal part recognition network using K positive sample images and L negative sample images until the comparison result between the output features of the second feature extraction network and the output features of the first feature extraction network reaches a convergence condition includes:

[0055] S1, use K positive sample images to train the first feature extraction network in the initialized abnormal part recognition network until the first sub-convergence condition is reached; S2, input K positive sample images and L negative sample images into the first feature extraction network that reaches the first sub-convergence condition and the second feature extraction network in the initialized abnormal part recognition network for training until the second sub-convergence condition is reached, wherein the positive sample image is input into the first feature extraction network that reaches the first sub-convergence condition to obtain the first sample sub-image feature, the positive sample image is input into the second feature extraction network to obtain the second sample sub-image feature, the negative sample image is input into the first feature extraction network that reaches the first sub-convergence condition to obtain the third sample sub-image feature, and the negative sample image is input into the second feature extraction network to obtain the fourth sample sub-image feature, and the second sub-convergence condition indicates that the feature similarity between the first sample sub-image feature and the second sample sub-image feature is greater than the seventh threshold, but the feature similarity between the third sample sub-image feature and the fourth sample sub-image feature is less than the eighth threshold.

[0056] It should be noted that the above-mentioned first sub-convergence condition may include, but is not limited to: the similarity between the image features obtained after the positive sample image is input into the first feature extraction network and the actual image features of the positive sample image is greater than the ninth threshold.

[0057] For example, the training of the first feature extraction network in the abnormal part recognition network initialized with K positive sample images until the first sub-convergence condition is met may include, but is not limited to: sequentially using the i-th positive sample image among the K positive sample images as the current positive sample image, and performing the following steps until the first sub-convergence condition is met: inputting the current positive sample image into the first feature extraction network to obtain image features output by the first feature extraction network; obtaining the similarity between the image features output by the first feature extraction network and the actual image features of the positive sample image; determining whether the similarity is greater than a ninth threshold, and if so, stopping training. If the similarity is less than the ninth threshold, obtaining the next positive sample image as the current positive sample image, and repeating the above steps.

[0058] It should be noted that the network structures used by the initialized first feature extraction network and the initialized second feature extraction network in this embodiment are the same, and may be, but are not limited to, a deep convolutional neural network structure (Residual Network-50, abbreviated as: ResNet-50), a convolutional neural network model (Vision Transformer, abbreviated as: ViT), a deep convolutional neural network model (Visual Geometry Group 16, abbreviated as: VGG-16) pre-trained model, etc., and are not limited to this in this embodiment.

[0059] As an optional embodiment, the aforementioned method of inputting K positive sample images and L negative sample images into the first feature extraction network that has reached the first sub-convergence condition and the second feature extraction network in the initialized abnormal part recognition network for training until the second sub-convergence condition is reached may include, but is not limited to: sequentially using the i-th positive sample image among the K positive sample images as the current positive sample image and the i-th negative sample image among the L negative sample images as the current negative sample image, and performing the following steps until the second sub-convergence condition is reached: inputting the current positive sample image and the current negative sample image, respectively, into the first feature extraction network that has reached the first sub-convergence condition to obtain first sample sub-image features and third sample sub-image features; and inputting the current positive sample image and the current negative sample image, respectively, into the second feature extraction network that has reached the second sub-convergence condition to obtain second sample sub-image features and fourth sample sub-image features. Then, determining whether the feature similarity between the first sample sub-image features and the second sample sub-image features is greater than a seventh threshold, and whether the feature similarity between the third sample sub-image features and the fourth sample sub-image features is less than an eighth threshold. Furthermore, when the feature similarity between the first sample sub-image feature and the second sample sub-image feature is greater than the seventh threshold, and the feature similarity between the third sample sub-image feature and the fourth sample sub-image feature is less than the eighth threshold, training is stopped.

[0060] S206, obtaining feature similarity between the first image feature and the second image feature;

[0061] S208 : Determine a recognition result of the image to be analyzed based on the feature similarity, wherein the recognition result is used to indicate whether the part object to be detected in the image to be analyzed is an abnormal part object.

[0062] It should be noted that in this embodiment, the feature similarity between the first image feature and the second image feature can be determined by using, but is not limited to, a regression loss function. For example, the following mean square error (L2 loss) regression loss function is used to obtain the feature similarity between the first image feature and the second image feature:

[0063] The parameter n in the above function is the number of image feature dimensions corresponding to the image to be analyzed, and the parameter x in the above function is i is the first image feature under the i-th dimension corresponding to the image to be analyzed, and the parameter y in the above function is iis the second image feature under the i-th dimension corresponding to the image to be analyzed. If the value of Loss(x, y) is greater than a certain threshold (for example, 0.5), it means that the gap between the first image feature and the second image feature is large, and the feature similarity is small, and then the part object to be detected in the image to be analyzed is determined to be an abnormal part object. If the value of Loss(x, y) is less than or equal to 0.5, it means that the gap between the first image feature and the second image feature is small, and the feature similarity is large, and then the part object to be detected in the image to be analyzed is determined to be a normal part object. It should be noted that other loss functions can also be used to determine the feature similarity between the first image feature and the second image feature, and this is not limited in this embodiment.

[0064] As an optional implementation, assume that the above image recognition method is applied to an anomaly detection scenario for industrial parts, assume that the part to be detected is a screw part, and the image to be analyzed is a captured image including the screw part. The following steps as shown in Figure 3 are used to illustrate the above method:

[0065] Execute S302 to obtain an image to be analyzed obtained by capturing an image of the screw part.

[0066] Then, S304 is executed to extract first and second image features of the image to be analyzed using the abnormal part recognition network 302 having a twin network structure. Specifically, the teacher network 304 in the abnormal part recognition network extracts the first image features from the image to be analyzed, and the student network 306 in the abnormal part recognition network extracts the second image features from the image to be analyzed.

[0067] Then, step S306 is performed to compare the first image feature with the second image feature to obtain a feature similarity between the first image feature and the second image feature.

[0068] Next, S308 is executed to determine whether the similarity between the first image feature and the second image feature is greater than a predetermined threshold. If so, S310-1 is executed to determine that the image to be analyzed belongs to the second type of image, indicating a normal part object, and that the screw part does not have a defect. If so, S310-2 is executed to determine that the image to be analyzed belongs to the first type of image, indicating an abnormal part object, and that the screw part does have a defect.

[0069] Among them, the image features extracted by the student network for the positive sample image are the same as the image features extracted by the first feature extraction network for the positive sample image, and the image features extracted by the student network for the negative sample image are different from the image features extracted by the teacher network for the negative sample image.

[0070] In an embodiment of the present application, an image to be analyzed is obtained by capturing an image of a part to be inspected. Then, an abnormal part recognition network having a Siamese network structure is used to extract image features from the image to be analyzed. The image features include first image features extracted by a first feature extraction network within the Siamese network structure and second image features extracted by a second feature extraction network within the Siamese network structure. The first feature extraction network is trained using positive sample images, while the second feature extraction network is trained using negative sample images, using the first feature extraction network as a reference network. Positive sample images are sample images from a sample image pair that include a normal part object, while negative sample images are sample images synthesized based on the positive sample images and include an abnormal part object. During the training process, the output of the second feature extraction network is aligned as closely as possible with the output of the first feature extraction network. This ensures that the second feature extraction network is robust to abnormal defects, while the first feature extraction network, which serves as the reference network, is not robust to abnormal defects. Therefore, during recognition, whether the part to be inspected is an abnormal part object can be determined by determining whether the features output by the first and second feature extraction networks are similar. Specifically, the feature similarity between the first and second image features is determined. Furthermore, the recognition result of the image to be analyzed is determined based on the feature similarity, wherein the recognition result is used to indicate whether the part object to be detected in the image to be analyzed is an abnormal part object. That is to say, in the embodiment of the present application, the similarity between the features of the image to be analyzed corresponding to the part object to be detected extracted by each branch network (i.e., the first feature extraction network, the second feature extraction network) in the abnormal part recognition network is used to determine whether the part object to be detected has abnormal defects. It is not a single determination of the recognition result of the image to be analyzed based on the similarity between the image features of the image corresponding to the part object to be detected and the image features of the image corresponding to the normal part object. This solves the technical problem of low accuracy of the recognition results of the image recognition method for industrial parts in the prior art, and achieves the technical effect of improving the accuracy of the recognition results of the image recognition method for industrial parts.

[0071] As an optional solution, determining the recognition result of the image to be analyzed based on feature similarity includes:

[0072] When the feature similarity is less than or equal to a first threshold, determining that the part object to be detected is an abnormal part object, and determining that the recognition result is that the image to be analyzed belongs to a first type of image for indicating an abnormal part object;

[0073] When the feature similarity is greater than or equal to a second threshold, determining that the part object to be detected is a normal part object, and determining that the recognition result is that the image to be analyzed belongs to a second type of image indicating a normal part object;

[0074] The first threshold is smaller than the second threshold.

[0075] Optionally, in this embodiment, the feature similarity can be obtained by using, but not limited to, a loss function between the first image feature and the second image feature. The larger the loss function value between the first image feature and the second image feature, the smaller the feature similarity between the first image feature and the second image feature; correspondingly, the smaller the loss function value between the first image feature and the second image feature, the greater the feature similarity between the first image feature and the second image feature.

[0076] As an optional implementation, assuming the similarity between the first image feature and the second image feature is 0.2, assuming the first threshold is 0.3, and the second threshold is 0.5, the above method is illustrated by the following steps: the first image feature is compared with the second image feature, and the feature similarity between the first image feature and the second image feature is 0.2. It is determined that the feature similarity between the first image feature and the second image feature (0.2) is less than the first threshold (0.3), and then the recognition result is determined to be that the image to be analyzed belongs to the first type of image indicating an abnormal part object. In other words, the part object to be detected is an abnormal part object.

[0077] As another optional implementation, assuming that the similarity between the first image feature and the second image feature is 0.6, and assuming that the first threshold is 0.3 and the second threshold is 0.5, the above method is illustrated by the following steps: the first image feature is compared with the second image feature, and the feature similarity between the first image feature and the second image feature is 0.6. It is determined that the feature similarity between the first image feature and the second image feature (0.6) is greater than the second threshold (0.5), and the recognition result is determined to be that the image to be analyzed belongs to the second type of image indicating a normal part object. In other words, the part object to be detected is a normal part object.

[0078] In an embodiment of the present application, when the feature similarity is less than or equal to a first threshold, the part to be detected is determined to be an abnormal part object, and the recognition result is that the image to be analyzed belongs to the first category of images indicating abnormal part objects. When the feature similarity is greater than or equal to a second threshold, the part to be detected is determined to be a normal part object, and the recognition result is that the image to be analyzed belongs to the second category of images indicating normal part objects. The first threshold is less than the second threshold. In other words, in this embodiment, the presence of an abnormal defect in the part to be detected is determined based on the similarity between the features of the image to be analyzed corresponding to the part to be detected, extracted by each branch network in the abnormal part recognition network (i.e., the first feature extraction network and the second feature extraction network). This is rather than simply determining the recognition result of the image to be analyzed based on the similarity between the image features of the image corresponding to the part to be detected and the image features of the image corresponding to a normal part object. This solves the technical problem of low recognition accuracy in existing industrial part image recognition methods, achieving the technical effect of improving the recognition accuracy of industrial part image recognition methods.

[0079] Optionally, as an optional solution, after acquiring the image to be analyzed obtained by performing image acquisition on the part object to be inspected, the above method further includes:

[0080] S1, obtaining a template image corresponding to a part object to be detected, wherein the template image includes a reference part object that belongs to the same type of part as the part object to be detected, and the reference part object is a normal part object;

[0081] Optionally, in this embodiment, the above-mentioned template image can be obtained based on but not limited to the following steps: obtaining an absolutely normal and flawless part object to be inspected as a reference part object, and then performing image acquisition on all image points that need to be inspected of the reference part object to obtain a template image.

[0082] S2, correcting the display position of the part object to be detected in the image to be analyzed according to the display position of the reference part object in the template image to obtain a corrected image, wherein the part object to be detected in the corrected image and the reference part object in the template image are in a display aligned state.

[0083] It should be noted that the display position of the part object to be detected in the image to be analyzed may include, but is not limited to, the display direction, angle, and position of the part object to be detected in the image to be analyzed. The display alignment of the part object to be detected with the reference part object in the template image may be, but is not limited to, indicating that the display direction, angle, and position of the part object to be detected in the image to be analyzed are consistent with the display direction, angle, and position of the reference part object in the template image. Furthermore, the display position of the part object to be detected in the image to be analyzed may be corrected using, but is not limited to, a transformation matrix between the image to be analyzed and the template image.

[0084] For example, assuming that the image shown in FIG4(a) is the image to be analyzed, the part object to be detected included in the image to be analyzed is a screw part, and the image shown in FIG4(b) is the template image, and the reference part object included in the template image is also a screw part. The above-mentioned correction of the display position of the part object to be detected in the image to be analyzed according to the display position of the reference part object in the template image can include, but is not limited to, adjusting the part object to be detected in FIG4(a) so that the display direction, angle, and position of the part object to be detected are consistent with the display direction, angle, and position of the reference part object in FIG4(b).

[0085] In an embodiment of the present application, a template image corresponding to a part object to be detected is obtained, wherein the template image displays a reference part object that belongs to the same type of part as the part object to be detected, and the reference part object is a normal part object. Then, according to the display position of the reference part object in the template image, the display position of the part object to be detected in the image to be analyzed is corrected to obtain a corrected image, wherein the part object to be detected in the corrected image is in a display alignment state with the reference part object in the template image. In other words, using an embodiment of the present application, a template image corresponding to the part object to be detected is used to correct the display position of the part object to be detected in the image to be analyzed, so that the image to be analyzed is more standardized, and further, the image features of the image to be analyzed extracted using the abnormal part recognition network are more accurate, thereby achieving the technical effect of improving the accuracy of the recognition results of the image recognition method for industrial parts.

[0086] Optionally, as an optional solution, the display position of the part object to be detected in the image to be analyzed is corrected according to the display position of the reference part object in the template image, and the corrected image is obtained, including:

[0087] S1, performing edge detection processing on the reference part object in the template image to obtain a first object contour map, and performing edge detection processing on the part object to be detected in the image to be analyzed to obtain a second object contour map;

[0088] It should be noted that edge detection processing can be performed on the reference part object in the template image and the part object to be detected in the image to be analyzed using, but not limited to, a differential method, a differential edge detection method, a Roberts edge detection operator, a Sobel edge detection operator, a Laplace edge detection operator, and the like. This is not limited in the present embodiment.

[0089] Among them, the Roberts edge detection operator uses the difference between two adjacent pixels in the diagonal direction based on the principle that the difference in any pair of mutually perpendicular directions can be used to calculate the gradient. Then the Roberts gradient amplitude value is calculated. The Roberts detector is relatively simple, but has some functional limitations. For example, it is asymmetric and cannot detect edges such as multiples of 45°. However, it is still often used in hardware implementation because it is simple and fast. The Sobel edge detection operator examines the weighted difference in the grayscale of each pixel of the digital image with its upper, lower, left and right neighbors. The Laplace edge detection operator is a second-order differential operator. Unlike other edge detection methods, this method is an isotropic detection method, that is, the degree of edge enhancement is independent of the direction of the edge, so it can meet the requirements of edge sharpening in different directions.

[0090] S2, converting the first object contour map into a first object contour point set, and converting the second object contour map into a second object contour point set;

[0091] Optionally, the above-mentioned conversion of the first object contour map into the first object contour point set may include, but is not limited to, the following: converting all key positions in the first object contour map into points in the first coordinate system using the non-maximum suppression (NMS) algorithm to obtain the first object contour point set. Correspondingly, the above-mentioned conversion of the second object contour map into the second object contour point set may include, but is not limited to, converting all key points in the second object contour map into points in the second coordinate system using the non-maximum suppression (NMS) algorithm to obtain the second object contour point set. Among them, the NMS algorithm is widely used in traditional feature extraction and deep learning target detection algorithms. The principle of the NMS algorithm is to obtain the optimal solution by screening out local maxima. In 2D edge extraction, it is reflected in the fact that after extracting the edge contour, some points with a small gradient direction change rate are screened out to avoid interference. It also plays an important role in 3D key point detection, screening out non-local extreme values ​​in features.

[0092] It should be noted that after converting the first object contour map into the first object contour point set and converting the second object contour map into the second object contour point set, the following steps may also be performed, but are not limited to: eliminating error points included in the first object contour point set and the second object contour point set.

[0093] For example, assuming that the image shown in (a) of FIG. 5 is the image to be analyzed, the part object to be detected included in the image to be analyzed is a screw part, and the image shown in (b) of FIG. 5 is the template image, and the reference part object included in the template image is also a screw part. The above-mentioned edge detection processing of the reference part object in the template image to obtain a first object contour map, and edge detection processing of the part object to be detected in the image to be analyzed to obtain a second object contour map may include, but is not limited to: performing edge detection processing on the part object to be detected in (a) of FIG. 5 to obtain a contour map of the part object to be detected as shown in (c) of FIG. 5 (i.e., a first object contour map). Performing edge detection processing on the reference part object in (b) of FIG. 5 to obtain a contour map of the reference part object as shown in (d) of FIG. 5 (i.e., a second object contour map).

[0094] Next, the above-mentioned conversion of the first object contour map into a first object contour point set and the conversion of the second object contour map into a second object contour point set may include, but is not limited to: using the NMS algorithm to convert the first object contour map of the reference part object in (c) of Figure 5 into a contour point set of the reference part object (i.e., the first object contour point set) as shown in (e) of Figure 5; and using the NMS algorithm to convert the second object contour map of the part object to be detected in (d) of Figure 5 into a contour point set of the part object to be detected (i.e., the second object contour point set) as shown in (f) of Figure 5.

[0095] Then, the error points that do not belong to the reference part object are removed from the outline point set of the reference part object (i.e., the first object outline point set) as shown in Figure 5(e), and the error points that do not belong to the part object to be detected are removed from the outline point set of the part object to be detected (i.e., the second object outline point set) as shown in Figure 5(f). The first object outline point set without error points, as shown in Figure 5(g), and the second object outline point set without error points, as shown in Figure 5(h), are obtained.

[0096] S3, determining a correction transformation matrix based on a point position correspondence relationship between the first object contour point set and the second object contour point set;

[0097] It should be noted that before determining the correction transformation matrix based on the point correspondence between the first object contour point set and the second object contour point set, the above method may, but is not limited to, also include: utilizing a nearest neighbor search method to determine the matching relationship between each point included in the first object contour point set and each point included in the second object contour point set. Nearest neighbor search, also known as closest point search, refers to the optimization problem of searching for the closest point to a query point in a scale space. Nearest neighbor search has a wide range of applications in many fields, such as computer vision, information retrieval, data mining, machine learning, and large-scale learning. It is most widely used in the field of computer vision, such as computer graphics, image retrieval, duplicate retrieval, object recognition, scene recognition, scene classification, posture assessment, and feature matching.

[0098] Alternatively, assume that the point set shown in FIG6(a) is the first object contour point set, and the point set shown in FIG6(b) is the second object contour point set. Taking point A in the first object contour point set and point A1 in the second object contour point set as examples, the above method is illustrated by the following steps: Determine the coordinates of each point in the first object contour point set in the first coordinate system of the first object contour point set, and determine the coordinates of each point in the second object contour point set in the second coordinate system of the second object contour point set. As shown in FIG6(a), the coordinates of point A in the first object contour point set in the first coordinate system are (-9, -2), and as shown in FIG6(b), the coordinates of point A1 in the second object contour point set in the second coordinate system are (-6, -8). Then, the nearest neighbor search method is used to determine the matching relationship between each point in the first object contour point set and each point in the second object contour point set. Among them, point A in the first object contour point set and point A1 in the second object contour point set are both points at the tail of the screw part, and the two have a matching relationship {(-9, -2), (-6, -8)}.

[0099] Furthermore, the above-mentioned determination of the correction transformation matrix based on the point correspondence between the first object contour point set and the second object contour point set may include, but is not limited to, including: determining the correction transformation matrix based on a machine vision (Random Sample Consensus, abbreviated as: ransac) algorithm and utilizing the point correspondence between the first object contour point set and the second object contour point set.

[0100] S4, using the correction transformation matrix to perform position correction on the second object contour point set in the image to be analyzed, to obtain a corrected image containing the third object contour point set, wherein the display position of each contour point in the third object contour point set in the corrected image will correspond to the display position of each contour point in the first object contour point set in the template image.

[0101] Optionally, in this embodiment, the above-mentioned use of the correction transformation matrix to perform position correction on the second object contour point set in the image to be analyzed to obtain a corrected image containing the third object contour point set may include, but is not limited to: multiplying the second object contour point set in the image to be analyzed by the correction transformation matrix to obtain a corrected image containing the third object contour point set.

[0102] In an embodiment of the present application, edge detection is performed on a reference part object in a template image to obtain a first object contour map. Edge detection is also performed on the part object to be detected in the image to be analyzed to obtain a second object contour map. The first object contour map is then converted into a first object contour point set, and the second object contour map is converted into a second object contour point set. A correction transformation matrix is ​​then determined based on the point-to-point correspondence between the first object contour point set and the second object contour point set. The correction transformation matrix is ​​then used to positionally correct the second object contour point set in the image to be analyzed, resulting in a corrected image containing a third object contour point set. The positions of the contour points in the third object contour point set in the corrected image correspond to the positions of the contour points in the first object contour point set in the template image. In other words, in this embodiment of the present application, the position of the part object to be detected in the image to be analyzed is corrected using the template image corresponding to the part object to be detected, resulting in a more standardized image to be analyzed. This, in turn, makes the image features extracted from the image to be analyzed using the abnormal part recognition network more accurate, thereby achieving the technical effect of improving the accuracy of recognition results for industrial part image recognition.

[0103] Optionally, as an optional solution, determining the correction transformation matrix based on the point position correspondence between the first object contour point set and the second object contour point set includes:

[0104] S1, determine the current transformation matrix to be estimated;

[0105] S2, transforming the second object contour point set using the current transformation matrix to obtain a reference contour point set;

[0106] S3, based on the point position correspondence between each contour point in the first object contour point set and each contour point in the reference contour point set, determining a position error between a displayed position of each contour point in the reference contour point set and a displayed position of each contour point in the first object contour point set, to obtain a plurality of point position errors;

[0107] It should be noted that, after determining the position errors between the display positions of each contour point in the reference contour point set and the display positions of each contour point in the first object contour point set based on the point position correspondence between each contour point in the first object contour point set and each contour point in the second object contour point set, and obtaining multiple point position errors, the above method further includes: eliminating contour points whose point position errors are greater than a fifth threshold from the reference contour point set, wherein the fifth threshold is an integer greater than 0.

[0108] S4, determining a transformation error between the reference contour point set and the first object contour point set using the plurality of point position errors;

[0109] S5, when the transformation error does not meet the error convergence condition, adjusting the current transformation matrix to obtain an adjusted transformation matrix, and using the adjusted transformation matrix to transform the second object contour point set;

[0110] S6. When the transformation error satisfies the error convergence condition, the current transformation matrix is ​​determined as the correction transformation matrix.

[0111] Optionally, in this embodiment, when the transformation error satisfies the error convergence condition, determining the current transformation matrix as the correction transformation matrix may be limited to including: when the transformation error is less than a sixth threshold value for N consecutive times, determining the current transformation matrix as the correction transformation matrix.

[0112] As an optional implementation manner, the above method is explained by taking the following steps as shown in FIG7 as an example:

[0113] Execute S702 to determine a current transformation matrix to be estimated, where the current transformation matrix is ​​a randomly initialized transformation matrix.

[0114] Then, step S704 is executed to transform the second object contour point set using the current transformation matrix to obtain a reference contour point set.

[0115] Then, S706 is executed to establish a point position correspondence relationship between each contour point in the reference contour point set and each contour point in the first object contour point set.

[0116] Next, execute S708 to determine the position errors between the display positions of each contour point in the reference contour point set and the display positions of each contour point in the first object contour point set based on the point position correspondence between each contour point in the first object contour point set and each contour point in the reference contour point set.

[0117] For example, assuming that the reference contour point set includes K contour points and the first object contour point set includes K contour points, then the position error between the kth contour point in the reference contour point set and the kth contour point in the first object contour point set is as shown in formula (2). It is the second-order Euclidean distance between the kth contour point in the reference contour point set and the kth contour point in the first object contour point set:

[0118] Among them, w k is the position error between the kth contour point in the reference contour point set and the kth contour point in the first object contour point set. k is the coordinate of the kth contour point in the reference contour point set in the second coordinate system, is the coordinate of the kth contour point in the first object contour point set in the first coordinate system.

[0119] The position error between the kth contour point in the reference contour point set and the kth contour point in the first object contour point set is:

[0120] Among them, W k is the position error between the kth contour point in the reference contour point set and the kth contour point in the first object contour point set. K is the total number of contour points included in the reference contour point set, which is also the total number of contour points included in the first object contour point set. k The position error between the kth contour point in the reference contour point set and the kth contour point in the first object contour point set.

[0121] Then, execute S710 to determine the transformation error between the reference contour point set and the first object contour point set using the multiple point position errors.

[0122] For example, assuming that the reference contour point set includes K contour points and the first object contour point set includes K contour points, the transformation error between the reference contour point set and the first object contour point set is:

[0123] Where F is the transformation error between the reference contour point set and the first object contour point set, W k is the position error between the kth contour point in the reference contour point set and the kth contour point in the first object contour point set.

[0124] Then, step S712 is executed to determine whether the transformation error satisfies an error convergence condition. For example, assuming that the transformation error between the reference contour point set and the first object contour point set is F, it is determined whether F is less than a sixth threshold.

[0125] If the transformation error satisfies the error convergence condition, S714 is executed to determine the current transformation matrix as the correction transformation matrix, and it is determined that the correction processing of the second object contour point set in the image to be analyzed is completed. If the transformation error does not satisfy the error convergence condition, S716 is executed to adjust the current transformation matrix. Specifically, based on the Ransac algorithm, an adjusted transformation matrix is ​​determined using the point position correspondence relationship between the first object contour point set and the reference contour point set, and the adjusted transformation matrix is ​​determined as the current transformation matrix.

[0126] Then, S718-S724 are executed to transform the reference contour point set using the adjusted current transformation matrix. A point-to-point correspondence relationship is established between each contour point in the reference contour point set and each contour point in the first object contour point set. Based on the point-to-point correspondence relationship between each contour point in the first object contour point set and each contour point in the reference contour point set, a positional error between the displayed position of each contour point in the reference contour point set and the displayed position of each contour point in the first object contour point set is determined. A transformation error between the reference contour point set and the first object contour point set is determined using the multiple point-to-point errors.

[0127] Furthermore, S726 is executed to determine whether the transformation error satisfies the error convergence condition. If the transformation error satisfies the error convergence condition, S728 is executed to determine the current transformation matrix as the correction transformation matrix, and to determine that correction processing of the second object contour point set in the image to be analyzed is complete. If the transformation error does not satisfy the error convergence condition, S716-S726 are repeatedly executed until the transformation error satisfies the error convergence condition.

[0128] In an embodiment of the present application, a current transformation matrix to be estimated is determined. The current transformation matrix is ​​then used to transform the second object contour point set to obtain a reference contour point set. Next, based on the point position correspondence between each contour point in the first object contour point set and each contour point in the second object contour point set, the positional errors between the displayed positions of each contour point in the reference contour point set and the displayed positions of each contour point in the first object contour point set are determined, resulting in a plurality of point position errors. Furthermore, the transformation error between the reference contour point set and the first object contour point set is determined using the plurality of point position errors. If the transformation error does not meet an error convergence condition, the current transformation matrix is ​​adjusted to obtain an adjusted transformation matrix, and the second object contour point set is transformed using the adjusted transformation matrix. If the transformation error meets the error convergence condition, the current transformation matrix is ​​determined as a correction transformation matrix. In other words, by using the embodiment of the present application, a template image corresponding to the part object to be detected is used to correct the display position of the part object to be detected in the image to be analyzed, so that the image to be analyzed is more standardized, and the image features of the image to be analyzed extracted by the abnormal part recognition network are more accurate, thereby achieving the technical effect of improving the accuracy of the recognition results of the image recognition method of industrial parts.

[0129] Optionally, as an optional solution, extracting image features of the image to be analyzed using an abnormal part recognition network having a twin network structure includes:

[0130] S1, dividing the image to be analyzed into regions to obtain N image blocks to be analyzed, where N is a positive integer greater than 1;

[0131] S2, using the abnormal parts recognition network to extract features from the N image blocks to be analyzed to obtain image features.

[0132] Optionally, in this embodiment, the above-mentioned regional division of the image to be analyzed to obtain N image blocks to be analyzed may include, but is not limited to, at least one of the following: dividing the image to be analyzed according to preset sizes to obtain N image blocks of the same size; determining the N image blocks as N image blocks to be analyzed; intercepting N key image blocks from the image to be analyzed, wherein the key image blocks are image blocks where key object parts of the part object to be detected are located in the image to be analyzed; and determining the N key image blocks as N image blocks to be analyzed.

[0133] Furthermore, the above-mentioned abnormal parts recognition network is used to extract features from the N image blocks to be analyzed, and the image features obtained may include but are not limited to: determining each of the N image blocks to be analyzed as the current image block in turn, and performing the following steps: inputting the current image block into a first feature extraction network to obtain a first current sub-image feature, and inputting the current image block into a second feature extraction network to obtain a second current sub-image feature, wherein the first image feature includes the first current sub-image feature, and the second image feature includes the second current sub-image feature.

[0134] In an embodiment of the present application, the image to be analyzed is divided into regions to obtain N image blocks to be analyzed, where N is a positive integer greater than 1. Then, the abnormal part recognition network is used to extract features from the N image blocks to be analyzed to obtain image features. In other words, according to an embodiment of the present application, by dividing the image to be analyzed into regions to obtain N target blocks, when the abnormal part recognition network is used to extract features from the N image blocks to be analyzed to obtain image features, as long as a problem is detected in one of the N image blocks to be analyzed, it can be determined that the part object to be detected is an abnormal part object, thereby achieving the technical effect of improving the recognition efficiency of image recognition. Furthermore, the image to be analyzed is divided into regions to obtain N target blocks, when the abnormal part recognition network is used to extract features from the N image blocks to be analyzed to obtain image features, when a problem is detected in one of the N image blocks to be analyzed, the abnormal position of the abnormal part object to be detected can be quickly located, thereby facilitating the repair of the abnormal part object to be detected.

[0135] Optionally, as an optional solution, dividing the image to be analyzed into regions to obtain N image blocks to be analyzed includes one of the following:

[0136] Dividing the image to be analyzed according to a preset size to obtain N image blocks of the same size; determining the N image blocks as N image blocks to be analyzed;

[0137] For example, assuming that the image to be analyzed is the image shown in FIG8(a), the image to be analyzed is divided according to a predetermined size as shown in FIG8(b), to obtain four target blocks of exactly the same size.

[0138] N key image blocks are intercepted from the image to be analyzed, wherein the key image blocks are image blocks where key object parts of the part to be detected are located in the image to be analyzed; and the N key image blocks are determined as N image blocks to be analyzed.

[0139] It should be noted that, in this embodiment, the key object portion of the part to be inspected can be used, but is not limited to, to indicate the portion of the part to be inspected that needs to be inspected for defects, such as the nut and thread on a screw part.

[0140] As an optional implementation, assuming the image to be analyzed is as shown in FIG9(a), the above method is illustrated by the following steps: As shown in FIG9(a), it is determined that the part to be inspected included in the image to be analyzed is a screw part, and it is determined that the nut and the end of the thread on the screw part need to be inspected. Then, as shown in FIG9(b), two key image blocks are intercepted from the image to be analyzed: the block containing the nut and the block containing the end of the thread, as shown in FIG9(b).

[0141] In an embodiment of the present application, the image to be analyzed is divided according to a preset size to obtain N image blocks of the same size; the N image blocks are determined as N image blocks to be analyzed. This allows, in the process of extracting features from the N image blocks to be analyzed using an abnormal part recognition network to obtain image features, to determine that the part object to be detected is an abnormal part object as long as a problem is detected in one of the N image blocks to be analyzed, thereby improving the recognition efficiency of image recognition. N key image blocks are intercepted from the image to be analyzed, wherein the key image blocks are image blocks containing the key object parts of the part object to be detected in the image to be analyzed; the N key image blocks are determined as N image blocks to be analyzed. This reduces the data required to be processed by the abnormal part recognition network, thereby further improving the recognition efficiency of image recognition.

[0142] Optionally, as an optional solution, an abnormal part recognition network is used to extract features from N image blocks to be analyzed, and the obtained image features include:

[0143] Each of the N image blocks to be analyzed is sequentially determined as a current image block, and the following steps are performed:

[0144] The current image block is input into a first feature extraction network to obtain a first current sub-image feature, and the current image block is input into a second feature extraction network to obtain a second current sub-image feature, wherein the first image feature includes the first current sub-image feature and the second image feature includes the second current sub-image feature.

[0145] Optionally, as an optional embodiment, it is assumed that the N image blocks to be analyzed include: image block A to be analyzed, image block B to be analyzed, and image block C to be analyzed. In the case where the part object to be detected in the image to be analyzed is a normal part object, the above method is explained by the following steps:

[0146] Image block A to be analyzed is input into the first feature extraction network to obtain first sub-image feature A1. Image block A to be analyzed is then input into the second feature extraction network to obtain second sub-image feature A2. Feature similarity is determined between first sub-image feature A1 and first sub-image feature A2, and the similarity is determined to be greater than a second threshold. Next, image block B to be analyzed is input into the first feature extraction network to obtain first sub-image feature B1. Image block B to be analyzed is then input into the second feature extraction network to obtain second sub-image feature B2. Feature similarity is determined between first sub-image feature B1 and first sub-image feature B2, and the similarity is determined to be greater than a second threshold. Next, image block C to be analyzed is input into the first feature extraction network to obtain first sub-image feature C1. Image block C to be analyzed is then input into the second feature extraction network to obtain second sub-image feature C2. Feature similarity is determined between first sub-image feature C1 and first sub-image feature C2, and the similarity is determined to be greater than a second threshold. This determines that the part object to be detected in the image to be analyzed is a normal part object.

[0147] Alternatively, as another optional embodiment, it is still assumed that the N image blocks to be analyzed include: image block A to be analyzed, image block B to be analyzed, and image block C to be analyzed. In the case where the part object to be detected in the image to be analyzed is an abnormal part object and the abnormal area is located at the portion corresponding to image block B to be analyzed, the above method is explained by the following steps:

[0148] Image block A to be analyzed is input into the first feature extraction network to obtain first sub-image feature A1. Image block A to be analyzed is then input into the second feature extraction network to obtain second sub-image feature A2. Feature similarity is determined between first sub-image feature A1 and first sub-image feature A2, and it is then determined whether the feature similarity is greater than a second threshold. Next, image block B to be analyzed is input into the first feature extraction network to obtain first sub-image feature B1. Image block B to be analyzed is then input into the second feature extraction network to obtain second sub-image feature B2. Feature similarity is determined between first sub-image feature B1 and first sub-image feature B2, and it is then determined whether the feature similarity is less than the first threshold, indicating that the part to be detected in the image to be analyzed is an abnormal part. Next, image block C to be analyzed is input into the first feature extraction network to obtain first sub-image feature C1. Image block C to be analyzed is then input into the second feature extraction network to obtain second sub-image feature C2. Feature similarity is determined between first sub-image feature C1 and first sub-image feature C2, and it is then determined whether the feature similarity is greater than a second threshold. It is then determined that the part object to be detected in the image to be analyzed is an abnormal part object, and the abnormal area of ​​the abnormal part object is located at the position corresponding to the image block B to be analyzed.

[0149] In an embodiment of the present application, each of the N image blocks to be analyzed is sequentially determined as a current image block, and the following steps are performed: the current image block is input into a first feature extraction network to obtain a first current sub-image feature, and the current image block is input into a second feature extraction network to obtain a second current sub-image feature, wherein the first image feature includes the first current sub-image feature, and the second image feature includes the second current sub-image feature. In other words, using an embodiment of the present application, by dividing the image to be analyzed into regions to obtain N target blocks, when the abnormal part recognition network is used to extract features from the N image blocks to be analyzed to obtain image features, as long as a problem is detected in one of the N image blocks to be analyzed, the part object to be detected can be determined to be an abnormal part object, thereby achieving the technical effect of improving the recognition efficiency of image recognition. Furthermore, the image to be analyzed is divided into regions to obtain N target blocks, so that in the process of extracting features from the N image blocks to be analyzed using the abnormal part recognition network to obtain image features, when a problem is detected in one of the N image blocks to be analyzed, the abnormal position in the abnormal part object to be detected can be quickly located, thereby facilitating the repair of the abnormal part object to be detected.

[0150] Optionally, as an optional solution, before acquiring the image to be analyzed obtained by performing image acquisition on the part object to be inspected, the above method includes:

[0151] S1, obtaining K positive sample images including normal part objects and K negative sample images including abnormal part objects, to obtain K sample image pairs, wherein the negative sample images are synthesized by adjusting the positive sample images, and K is a natural number greater than 1;

[0152] Optionally, in this embodiment, before obtaining K positive sample images including normal part objects and K negative sample images including abnormal part objects to obtain K sample image pairs, the method may further include, but is not limited to: sequentially correcting the display positions of the part objects included in each of the K positive sample images including normal part objects, and then synthesizing K negative sample images including abnormal part objects based on the K positive sample images including normal part objects. For specific methods of correcting the display positions, please refer to the embodiment described above regarding correcting the display positions of the part objects to be detected in the image to be analyzed. This will not be further described in this embodiment.

[0153] S2, using K sample images to train the initialized abnormal part recognition network until the comparison results between the output features of the second feature extraction network and the output features of the first feature extraction network reach a convergence condition.

[0154] Optionally, the aforementioned method of training the initialized abnormal part recognition network using K sample images until the comparison results between the output features of the second feature extraction network and the output features of the first feature extraction network reach a convergence condition may further include: training the initialized abnormal part recognition network using K positive sample images and K negative sample images. During the training process, the parameters of the first feature extraction network remain fixed, and the output features of the first feature extraction network for the K positive samples showing normal part objects are used to influence the training of the second feature extraction network, and the network parameters of the second feature extraction network are continuously adjusted. This is done so that the comparison results between the output features of the second feature extraction network and the output features of the first feature extraction network reach a convergence condition. Assuming that the parameters of the first feature extraction network remain fixed, the aforementioned method of achieving convergence between the output features of the second feature extraction network and the output features of the first feature extraction network may include, but is not limited to, the similarity between the first sample sub-image features obtained after the positive sample image is input into the first feature extraction network and the second sample sub-image features obtained after the positive sample image is input into the second feature extraction network being greater than a third threshold. However, the feature similarity between the third sample sub-image feature obtained after the negative sample image is input into the first feature extraction network and the fourth sample sub-image feature obtained after the negative sample image is input into the second feature extraction network is less than the fourth threshold.

[0155] In an embodiment of the present application, K positive sample images including normal part objects and K negative sample images including abnormal part objects are obtained to obtain K sample image pairs, wherein the negative sample images are synthesized by adjusting the positive sample images, and K is a natural number greater than 1. Then, the K sample image pairs are used to train the initialized abnormal part recognition network until the comparison results between the output features of the second feature extraction network and the output features of the first feature extraction network reach the convergence condition. In other words, using the embodiment of the present application, the abnormal part recognition network is trained with rich sample data, thereby improving the accuracy of the abnormal part recognition network. This achieves the technical effect of improving the accuracy of the recognition results of the industrial part image recognition method.

[0156] Optionally, as an optional solution, using K sample images to train the initialized abnormal part recognition network until the comparison result between the output feature of the second feature extraction network and the output feature of the first feature extraction network reaches a convergence condition includes:

[0157] Get the current sample image pair from the K sample image pairs and perform the following operations:

[0158] S1, inputting the current positive sample image in the current sample image pair into the first feature extraction network to obtain the current positive sample sub-image feature, and inputting the current negative sample image in the current sample image pair into the second feature extraction network to obtain the current negative sample sub-image feature;

[0159] S2, when the feature similarity between the current negative sample sub-image feature and the current positive sample sub-image feature is less than a third threshold, obtaining the next pair of sample images as the current sample image pair;

[0160] S3, when the feature similarity between the current negative sample sub-image feature and the current positive sample sub-image feature is greater than or equal to a third threshold, adding one to the training convergence statistical result;

[0161] S4: When the training convergence statistical result reaches a fourth threshold, determine that the convergence condition is met.

[0162] It should be noted that the network structures used by the initialized first feature extraction network and the initialized second feature extraction network in this embodiment are the same, and can be but not limited to a deep convolutional neural network structure (Residual Network-50, abbreviated as: ResNet-50), a convolutional neural network model (Vision Transformer, abbreviated as: ViT), a deep convolutional neural network model (Visual Geometry Group 16, abbreviated as: VGG-16) pre-training model, etc., and no limitation is imposed on this in this embodiment.

[0163] Optionally, the convergence statistics are used to count the number of times that the feature similarity between the current negative sample sub-image feature and the current positive sample sub-image feature is greater than or equal to a third threshold. If this number reaches a fourth threshold, the abnormal part recognition network is determined to be sufficiently reliable, and training is terminated.

[0164] Optionally, in this embodiment, before inputting the current positive sample image in the current sample image pair into the first feature extraction network and inputting the current negative sample image in the current sample image pair into the second feature extraction network, the method may include, but is not limited to: performing block division on the current positive sample image to obtain N current positive sample image blocks. Furthermore, performing block division on the current negative positive sample image to obtain N current negative sample image blocks. For specific block division methods, please refer to the above-described embodiments of performing block division on the image to be analyzed, which will not be further described in this embodiment.

[0165] Accordingly, the above-mentioned inputting the current positive sample image in the current sample image pair into the first feature extraction network and inputting the current negative sample image in the current sample image pair into the second feature extraction network may, but is not limited to, include: inputting N current positive sample image blocks corresponding to the current positive sample image into the first feature extraction network, and inputting N current negative sample image blocks corresponding to the current negative sample image into the second feature extraction network, where N is a positive integer.

[0166] Furthermore, in the case where the feature similarity between the current negative sample sub-image feature and the current positive sample sub-image feature is greater than or equal to the third threshold, the training convergence statistical result is added with one processing; in the case where the training convergence statistical result reaches the fourth threshold, determining that the convergence condition is met may include, but is not limited to: when the feature similarity between the image features obtained after each of the above-mentioned N current positive sample image blocks is input into the first feature extraction network and the image features obtained after each of the above-mentioned N current negative sample image blocks is input into the second feature extraction network is greater than or equal to the third threshold, the training convergence statistical result is added with one processing; in the case where the training convergence statistical result reaches the fourth threshold, determining that the convergence condition is met.

[0167] As an optional embodiment, the above method is explained by taking the following steps as shown in FIG10 as an example:

[0168] S1002, obtaining a positive sample image.

[0169] Specifically, K positive sample images including normal part objects are obtained.

[0170] S1004: Correct the positive sample image.

[0171] Specifically, the display position of the part object displayed in each of the K positive sample images is corrected.

[0172] S1006: Obtain sample image pairs.

[0173] Specifically, K positive sample images are adjusted in sequence to synthesize K negative sample images including abnormal part objects, thereby generating K sample image pairs, wherein each sample image pair includes a positive sample image and a negative sample image synthesized from the positive sample image.

[0174] S1008: Determine the current sample image pair.

[0175] Specifically, the i-th sample image pair among the K sample image pairs is determined as the current sample image pair, where i is a positive integer less than K.

[0176] S1010 , dividing the image in the sample image pair into blocks.

[0177] Specifically, the positive sample image in the current sample image pair is divided into N positive sample image blocks, and the negative sample image in the current sample image pair is also divided into N negative sample image blocks corresponding to each image block in the N positive sample image blocks in the positive sample image (for example, if the N positive sample image blocks in the positive sample image are thread blocks and nut blocks, then the N negative sample image blocks in the negative sample image should also be thread blocks and nut blocks).

[0178] S1012-1, extracting image features of the positive sample image based on the first feature extraction network.

[0179] Specifically, N positive sample image blocks are sequentially input into the first feature extraction network to obtain N positive sample image block features.

[0180] S1012-2, extracting image features of the negative sample image based on the second feature extraction network.

[0181] Specifically, N negative sample image blocks are sequentially input into the second feature extraction network to obtain N negative sample image block features.

[0182] S1014, feature comparison.

[0183] Specifically, the feature similarities between the N positive sample image block features and the N negative sample image block features are compared respectively.

[0184] If, after comparing the N positive sample image block features with the N negative sample image block features, it is determined that there are positive sample image block features and negative sample image block features whose similarity is less than a third threshold, S1016-1 is executed to adjust parameters in the second feature extraction network and obtain the next sample image pair as the current sample image pair.

[0185] If, after comparing the N positive sample image block features obtained after the N positive sample image blocks are input into the first feature extraction network with the N negative sample image block features obtained after the N negative sample image blocks are input into the second feature extraction network, it is determined that the similarities between the N positive sample image block features and the N negative sample image block features are both greater than or equal to a third threshold, then S1016-2 is executed to add one to the training convergence statistical result; and it is determined whether the training convergence statistical result reaches a fourth threshold. If so, training is determined to be complete; if not, the next sample image pair is obtained as the current sample image pair.

[0186] In an embodiment of the present application, a current sample image pair is obtained from K sample image pairs, and the following operations are performed: the current positive sample image in the current sample image pair is input into a first feature extraction network, and the current negative sample image in the current sample image pair is input into a second feature extraction network; if the feature similarity between the current negative sample sub-image feature and the current positive sample sub-image feature is less than a third threshold, the next sample image pair is obtained as the current sample image pair; if the feature similarity between the current negative sample sub-image feature and the current positive sample sub-image feature is greater than or equal to the third threshold, the training convergence statistics are incremented by one; and if the training convergence statistics reach a fourth threshold, convergence conditions are determined to have been met. This ensures that, if the part object is a normal part object, the image features extracted from the image corresponding to the part object using the first feature extraction network are identical to the image features extracted from the image corresponding to the part object using the second feature extraction network. If the part object is an abnormal part object, the image features extracted from the image corresponding to the part object using the first feature extraction network are different from the image features extracted from the image corresponding to the part object using the second feature extraction network. This ensures the accuracy of the recognition results of the image to be analyzed, thereby achieving the technical effect of improving the accuracy of the recognition results of the image recognition method for industrial parts.

[0187] Optionally, as an optional solution, K positive sample images showing normal part objects and K negative sample images showing abnormal part objects are obtained, and the obtained K sample image pairs include:

[0188] S1, obtain K positive sample images including normal part objects;

[0189] S2, adjusting K positive sample images to synthesize K negative sample images, wherein the positive sample image is a sample template image including a normal part object.

[0190] Optionally, in this embodiment, adjusting the K positive sample images to synthesize the K negative sample images may include, but is not limited to, at least one of the following:

[0191] 1) Add white noise to the positive sample image to generate the negative sample image;

[0192] 2) obtaining a plurality of abnormal part image blocks, wherein the abnormal part image blocks include part defects belonging to the same type of parts as the part object to be detected; overlaying one or at least two of the plurality of abnormal part image blocks on the positive sample image to generate a negative sample image;

[0193] 3) Inputting the positive sample image into the abnormal part generation network to obtain a negative sample image, wherein the abnormal part generation network is a network obtained after training using the positive sample image and abnormal part image blocks to generate images including abnormal part objects, and the abnormal part image blocks include part defects belonging to the same type of parts as the part object to be detected.

[0194] In this embodiment of the present application, K positive sample images containing normal part objects are obtained. These K positive sample images are then adjusted to synthesize K negative sample images, where the positive sample images are sample template images containing normal part objects. In other words, in this embodiment of the present application, a rich set of negative sample images is obtained using K positive sample images containing normal part objects, enabling the second feature extraction network trained on these rich negative sample images to have good robustness. This achieves the technical effect of improving the accuracy of recognition results for industrial part image recognition.

[0195] Optionally, as an optional solution, adjusting the K positive sample images to synthesize the K negative sample images includes one of the following:

[0196] S1, add white noise to the positive sample image to generate a negative sample image;

[0197] Optionally, in this embodiment, the adding of white noise to the positive sample image to generate the negative sample image may include, but is not limited to: adding Gaussian noise to the positive sample image to generate the negative sample image.

[0198] S2, obtaining a plurality of abnormal part image blocks, wherein the abnormal part image blocks include part defects belonging to the same type of parts as the part object to be detected; overlaying one or at least two of the plurality of abnormal part image blocks on the positive sample image to generate a negative sample image;

[0199] It should be noted that, after covering one or at least two of the multiple abnormal part image blocks on the positive sample image, the method further includes: performing harmonization processing on the positive sample image block covered with the abnormal part image block.

[0200] Furthermore, assuming that the positive sample image is an image corresponding to a screw part as shown in Figure 11 (a), the abnormal part image block can be, but is not limited to, used to indicate the defective image block corresponding to the location where defects often occur in the screw part, as shown in Figure 11 (b), that is, the image block corresponding to the defective nut. The positive sample image block covered with the abnormal part image block can be, but is not limited to, as shown in Figure 11 (c), where the image block corresponding to the defective nut is covered on the nut area of ​​the screw part in the positive sample image.

[0201] S3, inputting the positive sample image into the abnormal part generation network to obtain a negative sample image, wherein the abnormal part generation network is a network obtained after training using the positive sample image and abnormal part image blocks, and is used to generate images including abnormal part objects, and the abnormal part image blocks include part defects belonging to the same type of parts as the part object to be detected.

[0202] It should be noted that before inputting the positive sample image into the abnormal part generation network to obtain the negative sample image, the method may include, but is not limited to, obtaining a positive sample image and a sample abnormal part image block, continuously training the initialized abnormal part generation network using the positive sample image and the sample abnormal part image block, and adjusting the parameters of the abnormal part generation network. This is done until a positive sample image and an abnormal part image block are input into the abnormal part generation network, resulting in a negative sample image belonging to the same class as the part to be detected.

[0203] Furthermore, the above-mentioned inputting the positive sample image into the abnormal part generation network to obtain the negative sample image may include, but is not limited to: inputting the positive sample image into the abnormal part generation network and the abnormal part image block into the abnormal part generation network to obtain the negative sample image.

[0204] In this embodiment, white noise is added to a positive sample image to generate a negative sample image; multiple abnormal part image blocks are obtained, wherein the abnormal part image blocks include part defects belonging to the same class as the part object to be detected; one or at least two of the multiple abnormal part image blocks are overlaid on the positive sample image to generate a negative sample image; and the positive sample image is input into an abnormal part generation network to obtain a negative sample image. The abnormal part generation network is a network trained using positive sample images and abnormal part image blocks to generate images containing abnormal part objects, wherein the abnormal part image blocks include part defects belonging to the same class as the part object to be detected. In other words, in this embodiment of the application, a rich set of negative sample images is obtained by using K positive sample images showing normal part objects, so that the second feature extraction network trained with these rich negative sample images has good robustness. This achieves the technical effect of improving the recognition accuracy of industrial part image recognition methods.

[0205] As an optional implementation, the following steps shown in FIG12 are used to illustrate the training steps of the abnormal part recognition network:

[0206] S1202: Acquire a positive sample image.

[0207] S1204: Perform registration processing on the positive sample image using an image registration module.

[0208] S1206: Divide the positive sample image into blocks using an image block division module.

[0209] S1208-1, extracting image features of the positive sample image based on the teacher network in the learning module.

[0210] S1208-2: synthesize a negative sample image based on the positive sample image in the negative sample image block module.

[0211] S1208-3, extracting image features of positive sample images and image features of negative sample images based on the student network in the learning module.

[0212] S1210 , in a feature comparison module, perform feature comparison on the image features of the positive sample image extracted based on the teacher network, the image features of the positive sample image extracted based on the student network, and the image features of the negative sample image.

[0213] Specifically, if the loss function between the image features of the positive sample image extracted based on the teacher network and the image features of the positive sample image extracted based on the student network is less than a predetermined threshold, and the loss function between the image features of the positive sample image extracted based on the teacher network and the image features of the negative sample image extracted based on the student network is less than a predetermined threshold, the training is terminated. Otherwise, the parameters of the student network are adjusted, and S1202-S1210 are repeatedly performed until the loss function between the image features of the positive sample image extracted based on the teacher network and the image features of the positive sample image extracted based on the student network is less than a predetermined threshold, and the loss function between the image features of the positive sample image extracted based on the teacher network and the image features of the negative sample image extracted based on the student network is less than a predetermined threshold.

[0214] It should be noted that the teacher network is used to represent the first feature extraction network mentioned above, and the student network is used to represent the second feature extraction network mentioned above. In the embodiments of the present application, the abnormal part recognition network is trained with rich sample data, thereby improving the accuracy of the abnormal part recognition network. This achieves the technical effect of improving the accuracy of recognition results of industrial part image recognition methods.

[0215] As an optional implementation, the following steps shown in FIG13 are used to illustrate the application steps of the abnormal part recognition network:

[0216] S1302: Acquire an image to be analyzed.

[0217] S1304: The image registration module performs registration processing on the image to be analyzed.

[0218] S1306: Divide the image to be analyzed into blocks using an image block division module.

[0219] S1308-1, extracting image features of the image to be analyzed based on the teacher network in the learning module.

[0220] S1308-2, extracting image features of the image to be analyzed based on the student network in the learning module.

[0221] S1310 , in a feature comparison module, compare the image features of the image to be analyzed extracted based on the teacher network with the image features of the positive sample image extracted based on the student network to obtain feature similarity.

[0222] S1312: Determine a recognition result of the image to be analyzed based on the feature similarity.

[0223] Specifically, when the loss function between the image features of the image to be analyzed extracted based on the teacher network and the image features of the image to be analyzed extracted based on the student network is less than a predetermined threshold, the recognition result is determined to indicate that the part object to be detected is a normal part object; when the loss function between the image features of the image to be analyzed extracted based on the teacher network and the image features of the image to be analyzed extracted based on the student network is greater than or equal to a predetermined threshold, the recognition result is determined to indicate that the part object to be detected is an abnormal part object.

[0224] It should be noted that the above-mentioned teacher network is used to indicate the first feature extraction network mentioned above, and the above-mentioned student network is used to indicate the second feature extraction network mentioned above. In an embodiment of the present application, the similarity between the features of the image to be analyzed corresponding to the part object to be detected extracted by each branch network in the abnormal part recognition network (i.e., the first feature extraction network, the second feature extraction network) is used to determine whether the part object to be detected has abnormal defects. It is not a single determination of the recognition result of the image to be analyzed based on the similarity between the image features of the image corresponding to the part object to be detected and the image features of the image corresponding to the normal part object. Thereby solving the technical problem of low accuracy of the recognition result of the image recognition method of industrial parts in the prior art, and achieving the technical effect of improving the accuracy of the recognition result of the image recognition method of industrial parts.

[0225] As another optional implementation, the modules used in the above image recognition method include:

[0226] 1) Image registration module:

[0227] The image registration module first needs to collect a set of template images. The collection method is to use a golden sample (an absolutely normal sample that has been tested) and take pictures of the sample at all image points that need to be tested. The collected set of pictures is the template image. The purpose of image registration is to correct the input test image so that its position can be completely aligned with the template image of the corresponding point, so that the image blocks of the specific area can be accurately extracted later. Specifically, the position and angle of the image to be analyzed and the template image are quite different. Through the image registration module, the transformation matrix T between the template image and the image to be analyzed can be calculated. By applying the transformation matrix T, the registered image to be analyzed can be obtained. The visual position and angle of the registered image to be analyzed are aligned with the template image. Subsequently, the corresponding image blocks in the image to be analyzed can be intercepted according to the pre-calibrated detection coordinate frame in the template image.

[0228] In this embodiment, the image registration module algorithm uses the minimum error iteration method. An error function is constructed, parameters to be estimated are defined, and based on the current estimate, an iterative algorithm is used to optimize the parameters to gradually reduce the error function. Specifically, edge detection is performed on the image to be analyzed and the template image to obtain their corresponding contour maps. The contour maps are then partitioned and subjected to parallel NMS operations to convert them into 2D contour point sets. The average error of the 2D contour point sets between the image to be analyzed and the template image is the error function. The specific function is as follows:

[0229] Among them, F is the error function, K is the number of contour points in the template image, and p k is the coordinate of the kth contour point in the image to be analyzed, is the coordinate of the kth contour point in the template image. After setting the error function, the first step is to apply the current estimated transformation matrix (initially the identity matrix) to the contour point set to be registered. The second step is to use the nearest neighbor search to establish a matching relationship between the contour point set of the image to be analyzed and the template image. The third step is to eliminate matching pairs with excessive errors. The fourth step is to estimate the transformation matrix using RANSAC, and then repeat the first step until the error function converges.

[0230] 2) Image block division module:

[0231] One of the purposes of the image block division module is to reduce the receptive field of a single anomaly detection image. Specifically, since many defects are small in size and mild in degree, if the entire image is input into the anomaly detection network for detection, since most areas in the image are normal areas and the abnormal areas only account for a small part, the detection difficulty will be very high. However, if the image blocks are divided first and then detected, the abnormal areas are larger than the area of ​​the image blocks, and the detection difficulty will be reduced. The second purpose of the image block division module is to locate the position of the abnormal defect. If the entire image is input into the anomaly detection network for detection, even if a certain image can be determined to be an abnormal image, it is difficult to locate the position of the abnormal defect. However, if the image blocks are divided first and then detected, as long as a certain image block is determined to be abnormal, the entire image can be determined to be abnormal, and the position of the image block in the image is the location of the abnormal defect.

[0232] In this embodiment, an image can be directly divided into 28×28 blocks of equal size, but this method is not limited to this. This division method is highly universal and can be used for any anomaly detection task. Furthermore, the division regions can be manually set based on a template image. After the image to be analyzed is registered based on the template image, it can be divided according to the manually set division regions. This can ensure that the divided regions have more distinct semantic features.

[0233] 3) Abnormal image block synthesis module:

[0234] The purpose of the abnormal image block synthesis module is to improve the robustness of the student network to abnormal defects. In this embodiment, the abnormal image blocks can be synthesized by, but not limited to, the following methods to improve the defect diversity of the abnormal image blocks:

[0235] Method 1: Adding white noise. Specifically, randomly adding some white noise to a normal image block can generate a synthetic abnormal image block. The size and position of the noise can be set within a range, and different random white noises can be added to simulate defects of various degrees and sizes.

[0236] Method 2: Paste defective image blocks. Specifically, for certain existing defects, use image clipping software to cut them out and then randomly paste them into normal image blocks. Then, perform image harmonization on the composite image blocks to obtain usable abnormal image blocks.

[0237] Method 3: Generate defects through a network. Specifically, a lightweight image processing technology (Inpainting) network is designed. The Inpainting network is trained using a normal image of the defective part and the defective part. When the normal image of the defective part and the defective part are input into the Inpainting network, the output result is a synthetic abnormal image block with the defect.

[0238] 4) Feature learning module:

[0239] In this embodiment, the network structure and initial parameters of the student network and the teacher network are the same, and a pre-trained ResNet-50 network can be used. During the learning process, the parameters of the teacher network are fixed, while the parameters of the student network are learnable. Specifically, the same normal image block is input to the student network and the teacher network so that their output features are as similar as possible. The purpose is to make it so that during the test phase, for the same normal image block input, the student network and the teacher network will output similar features. In addition, while the normal image block is input to the teacher network, the abnormal image block synthesized with the normal image block is also input to the student network so that its output features are as similar as possible. The purpose is to make it so that during the test phase, for the same abnormal image block input, the teacher network will output the features of the abnormal image block, while the student network, due to its robustness to abnormal defects, will output the features of the normal image block, so that the teacher network and the student network will output features with lower similarity.

[0240] Furthermore, in this embodiment, the loss function of the feature learning module is the normalized L2 loss function, which ranges from 0 to 1. The more similar the two features are, the smaller the loss function is.

[0241] 5) Feature comparison module:

[0242] The feature comparison module is the judgment module in the testing phase. For the input image to be analyzed, the features output by the teacher network and the student network are calculated. The L2 loss between the two features is then calculated. If the loss value is less than 0.5, the similarity between the two features is high, indicating that the image to be analyzed is normal. If the loss value is greater than 0.5, the similarity between the two features is low, indicating that the image to be analyzed is abnormal. Furthermore, the image to be analyzed is abnormal, and the location of the abnormal defect is the location of the abnormal image block.

[0243] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.

[0244] According to another aspect of the embodiments of the present application, an image recognition device for implementing the above-mentioned image recognition method is also provided. As shown in FIG14 , the device includes:

[0245] The first acquisition unit 1402 is used to acquire an image to be analyzed obtained by collecting an image of the part to be inspected;

[0246] Identification unit 1404 is configured to extract image features of an image to be analyzed using an abnormal part identification network having a twin network structure, wherein the image features include first image features extracted from the image to be analyzed by a first feature extraction network in the twin network structure, and second image features extracted from the image to be analyzed by a second feature extraction network in the twin network structure, the first feature extraction network being trained using positive sample images, and the second feature extraction network being trained using negative sample images with the first feature extraction network as a reference network, the positive sample image being a sample image in a sample image pair including a normal part object, and the negative sample image being a sample image synthesized based on the positive sample image and including an abnormal part object;

[0247] The second acquiring unit 1406 is configured to acquire a feature similarity between the first image feature and the second image feature;

[0248] The determining unit 1408 is configured to determine a recognition result of the image to be analyzed based on the feature similarity, wherein the recognition result is used to indicate whether the part object to be detected in the image to be analyzed is an abnormal part object.

[0249] Optionally, the determining unit includes:

[0250] a first determining module configured to determine, when the feature similarity is less than or equal to a first threshold, that the part object to be detected is an abnormal part object, and to determine, as a result of the recognition, that the image to be analyzed belongs to a first type of image indicating an abnormal part object;

[0251] The second determination module is used to determine that the part object to be detected is a normal part object when the feature similarity is greater than or equal to a second threshold, and to determine that the recognition result is that the image to be analyzed belongs to a second type of image used to indicate a normal part object; wherein the first threshold is less than the second threshold.

[0252] Optionally, the above device further includes:

[0253] a third acquiring unit, configured to acquire a template image corresponding to the part object to be detected, wherein the template image includes a reference part object that belongs to the same type of part as the part object to be detected, and the reference part object is a normal part object;

[0254] The correction unit is used to correct the display position of the part object to be detected in the image to be analyzed according to the display position of the reference part object in the template image, so as to obtain a corrected image, wherein the part object to be detected in the corrected image and the reference part object in the template image are in a display aligned state.

[0255] Optionally, the correction unit includes:

[0256] a detection module configured to perform edge detection processing on the reference part object in the template image to obtain a first object contour map, and to perform edge detection processing on the part object to be detected in the image to be analyzed to obtain a second object contour map;

[0257] a conversion module, configured to convert the first object contour map into a first object contour point set, and convert the second object contour map into a second object contour point set;

[0258] A third determining module is configured to determine a correction transformation matrix based on a point position correspondence relationship between the first object contour point set and the second object contour point set;

[0259] The correction module is used to perform position correction on the second object contour point set in the image to be analyzed using a correction transformation matrix to obtain a corrected image including a third object contour point set, wherein the display position of each contour point in the third object contour point set in the corrected image will correspond to the display position of each contour point in the first object contour point set in the template image.

[0260] Optionally, the above-mentioned third determination module is also used to: determine the current transformation matrix to be estimated; use the current transformation matrix to transform the second object contour point set to obtain a reference contour point set; based on the point position correspondence between each contour point in the first object contour point set and each contour point in the reference contour point set, determine the position error between the display position of each contour point in the reference contour point set and the display position of each contour point in the first object contour point set to obtain multiple point position errors; use the multiple point position errors to determine the transformation error between the reference contour point set and the first object contour point set; when the transformation error meets the error convergence condition, determine the current transformation matrix as the correction transformation matrix.

[0261] Optionally, the identification unit includes:

[0262] A division module is used to divide the image to be analyzed into regions to obtain N image blocks to be analyzed, where N is a positive integer greater than 1;

[0263] The extraction module is used to use the abnormal part recognition network to extract features from N image blocks to be analyzed to obtain image features.

[0264] Optionally, the above-mentioned division module is also used to divide the image to be analyzed according to a preset size to obtain N image blocks of the same size; determine the N image blocks as N image blocks to be analyzed; intercept N key image blocks from the image to be analyzed, wherein the key image blocks are image blocks where the key object parts of the part object to be detected in the image to be analyzed are located; and determine the N key image blocks as N image blocks to be analyzed.

[0265] Optionally, the above-mentioned extraction module is also used to determine each of the N image blocks to be analyzed as the current image block in turn, and perform the following steps: input the current image block into the first feature extraction network to obtain a first current sub-image feature, and input the current image block into the second feature extraction network to obtain a second current sub-image feature, wherein the first image feature includes the first current sub-image feature, and the second image feature includes the second current sub-image feature.

[0266] Optionally, the above device further includes:

[0267] a fourth acquisition unit, configured to acquire K positive sample images including normal part objects and K negative sample images including abnormal part objects, to obtain K sample image pairs, wherein the negative sample images are synthesized by adjusting the positive sample images, and K is a natural number greater than 1;

[0268] The training unit is used to train the initialized abnormal part recognition network using K sample images until the comparison result between the output feature of the second feature extraction network and the output feature of the first feature extraction network reaches a convergence condition.

[0269] Optionally, the above-mentioned training unit is also used to: obtain a current sample image pair from K sample image pairs, and perform the following operations: input the current positive sample image in the current sample image pair into the first feature extraction network to obtain the current positive sample sub-image feature, and input the current negative sample image in the current sample image pair into the second feature extraction network to obtain the current negative sample sub-image feature; when the feature similarity between the current negative sample sub-image feature and the current positive sample sub-image feature is less than a third threshold, obtain the next pair of sample image pairs as the current sample image pair; when the feature similarity between the current negative sample sub-image feature and the current positive sample sub-image feature is greater than or equal to the third threshold, add one to the training convergence statistical result; when the training convergence statistical result reaches a fourth threshold, determine that the convergence condition is met.

[0270] Optionally, the fourth obtaining unit includes:

[0271] An acquisition module is used to acquire K positive sample images including normal part objects;

[0272] The adjustment module is used to adjust K positive sample images to synthesize K negative sample images, wherein the positive sample images are sample template images including normal part objects.

[0273] Optionally, the adjustment module is also used to: add white noise to the positive sample image to generate a negative sample image; obtain multiple abnormal part image blocks, wherein the abnormal part image blocks include part defects that belong to the same type of parts as the part object to be detected; overlay one or at least two abnormal part image blocks among the multiple abnormal part image blocks on the positive sample image to generate a negative sample image; input the positive sample image into the abnormal part generation network to obtain a negative sample image, wherein the abnormal part generation network is a network obtained by training with positive sample images and abnormal part image blocks and is used to generate images including abnormal part objects, and the abnormal part image blocks include part defects that belong to the same type of parts as the part object to be detected.

[0274] For specific embodiments, please refer to the embodiments of the above-mentioned image recognition method, which will not be described in detail here.

[0275] According to another aspect of the embodiments of the present application, an electronic device for implementing the above-mentioned image recognition method is also provided. This embodiment is described using the electronic device as a terminal as an example. As shown in Figure 15, the electronic device includes a memory 1502 and a processor 1504. The memory 1502 stores a computer program, and the processor 1504 is configured to execute the steps of any of the above-mentioned method embodiments through the computer program.

[0276] Optionally, in this embodiment, the electronic device may be located in at least one network device among a plurality of network devices of a computer network.

[0277] Optionally, in this embodiment, the processor may be configured to execute any image recognition method of the aforementioned embodiments through a computer program.

[0278] Alternatively, those skilled in the art will appreciate that the structure shown in FIG15 is for illustration only, and the electronic device may also be a smartphone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile internet device (MID), a PAD, or other terminal device. FIG15 does not limit the structure of the electronic device. For example, the electronic device may include more or fewer components (such as a network interface, etc.) than those shown in FIG15 , or may have a configuration different from that shown in FIG15 .

[0279] Memory 1502 can be used to store software programs and modules, such as program instructions / modules corresponding to the image recognition method and apparatus in the embodiments of the present application. Processor 1504 executes the software programs and modules stored in memory 1502 to perform various functional applications and data processing, thereby implementing the aforementioned image recognition method. Memory 1502 can include high-speed random access memory (RAM) and can also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some embodiments, memory 1502 can further include memory remotely located from processor 1504, which can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof. Memory 1502 can specifically, but is not limited to, be used to store images to be analyzed. As an example, as shown in FIG. 15 , memory 1502 can include, but is not limited to, the first acquisition unit 1402, recognition unit 1404, second acquisition unit 1406, and determination unit 1408 of the aforementioned image recognition apparatus. In addition, it may also include but is not limited to other module units in the above-mentioned image recognition device, which will not be repeated in this example.

[0280] Optionally, the transmission device 1506 is used to receive or send data via a network. Specific examples of the network may include a wired network and a wireless network. In one embodiment, the transmission device 1506 includes a network interface controller (NIC), which can be connected to other network devices and a router via a network cable to communicate with the Internet or a local area network. In one embodiment, the transmission device 1506 is a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0281] In addition, the electronic device further includes: a connection bus 1508 for connecting various module components in the electronic device.

[0282] In other embodiments, the terminal device or server may be a node in a distributed system, wherein the distributed system may be a blockchain system, and the blockchain system may be a distributed system formed by connecting multiple nodes via network communication. The nodes may form a point-to-point network, and any computing device, such as a server, terminal, or other electronic device, may become a node in the blockchain system by joining the point-to-point network.

[0283] According to one aspect of the present application, a computer program product is provided, comprising a computer program containing program code for executing the above-described method. In such an embodiment, the computer program can be downloaded and installed from a network via a communication component and / or installed from a removable medium. When the computer program is executed by a central processing unit, the various functions provided in the embodiments of the present application are performed.

[0284] According to one aspect of the present application, a computer-readable storage medium is provided. A processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the above method.

[0285] Optionally, in this embodiment, the computer-readable storage medium may be configured to store a computer program for executing any one of the image recognition methods of the aforementioned embodiments.

[0286] Optionally, in this embodiment, a person of ordinary skill in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by instructing the hardware related to the terminal device through a program, and the program may be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0287] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above-mentioned computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling one or more computer devices (which can be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application.

[0288] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0289] In the several embodiments provided in this application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, and can be electrical or other forms.

[0290] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0291] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0292] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. An image recognition method, the method being performed by an electronic device, comprising: Acquire an image to be analyzed obtained by collecting an image of the part to be inspected; The image features of the image to be analyzed are extracted using an abnormal part recognition network having a twin network structure, wherein the image features include a first image feature extracted from the image to be analyzed by a first feature extraction network in the twin network structure, and a second image feature extracted from the image to be analyzed by a second feature extraction network in the twin network structure, the first feature extraction network is trained using a positive sample image, the second feature extraction network is trained using a negative sample image with the first feature extraction network as a reference network, the positive sample image is a sample image including a normal part object in a sample image pair, and the negative sample image is a sample image including an abnormal part object synthesized based on the positive sample image; Acquire feature similarity between the first image feature and the second image feature; A recognition result of the image to be analyzed is determined based on the feature similarity, wherein the recognition result is used to indicate whether the part object to be detected in the image to be analyzed is an abnormal part object.

2. According to the method of claim 1, determining the recognition result of the image to be analyzed based on the feature similarity comprises: In the case where the feature similarity is less than or equal to a first threshold, determining that the part object to be detected is the abnormal part object, and determining that the recognition result is that the image to be analyzed belongs to a first type of image for indicating an abnormal part object; In the case where the feature similarity is greater than or equal to a second threshold, determining that the part object to be detected is the normal part object, and determining that the recognition result is that the image to be analyzed belongs to a second type of image indicating a normal part object; The first threshold is smaller than the second threshold.

3. The method according to claim 1 or 2, after acquiring the image to be analyzed obtained by performing image acquisition on the part object to be inspected, the method further comprises: Acquire a template image corresponding to the part object to be detected, wherein the template image includes a reference part object that belongs to the same type of part as the part object to be detected, and the reference part object is a normal part object; According to the display position of the reference part object in the template image, the display position of the part object to be detected in the image to be analyzed is corrected to obtain a corrected image, wherein the part object to be detected in the corrected image and the reference part object in the template image are in a display alignment state.

4. The method according to claim 3, wherein the display position of the part object to be detected in the image to be analyzed is corrected according to the display position of the reference part object in the template image to obtain the corrected image, comprising: Performing edge detection processing on the reference part object in the template image to obtain a first object contour map, and performing edge detection processing on the part object to be detected in the image to be analyzed to obtain a second object contour map; converting the first object contour map into a first object contour point set, and converting the second object contour map into a second object contour point set; Determining a correction transformation matrix based on a point position correspondence relationship between the first object contour point set and the second object contour point set; The correction transformation matrix is ​​used to perform position correction on the second object contour point set in the image to be analyzed to obtain the corrected image containing the third object contour point set, wherein the display position of each contour point in the third object contour point set in the corrected image will correspond to the display position of each contour point in the first object contour point set in the template image.

5. The method according to claim 4, wherein determining the correction transformation matrix based on the point correspondence relationship between the first object contour point set and the second object contour point set comprises: determining a current transformation matrix to be estimated; Using the current transformation matrix to transform the second object contour point set to obtain a reference contour point set; Based on the point position correspondence between each contour point in the first object contour point set and each contour point in the reference contour point set, determining a position error between each display position of each contour point in the reference contour point set and each contour point in the first object contour point set to obtain a plurality of point position errors; Determine a transformation error between the reference contour point set and the first object contour point set using the multiple point position errors; In a case where the transformation error satisfies the error convergence condition, the current transformation matrix is ​​determined as the correction transformation matrix.

6. According to the method of any one of claims 1 to 5, the extracting the image features of the image to be analyzed by using an abnormal part recognition network having a twin network structure comprises: Dividing the image to be analyzed into regions to obtain N image blocks to be analyzed, where N is a positive integer greater than 1; The abnormal part recognition network is used to extract features from the N image blocks to be analyzed to obtain the image features.

7. According to the method of claim 6, the step of dividing the image to be analyzed into regions to obtain N image blocks to be analyzed comprises one of the following: Dividing the image to be analyzed according to a preset size to obtain N image blocks with the same size; determining the N image blocks as the N image blocks to be analyzed; N key image blocks are intercepted from the image to be analyzed, where: The key image block is an image block where the key object portion of the part object to be detected in the image to be analyzed is located; and the N key image blocks are determined as the N image blocks to be analyzed.

8. According to the method of claim 6 or 7, the step of extracting features from the N image blocks to be analyzed using the abnormal parts recognition network to obtain the image features comprises: Each of the N image blocks to be analyzed is sequentially determined as a current image block, and the following steps are performed: The current image block is input into the first feature extraction network to obtain a first current sub-image feature, and the current image block is input into the second feature extraction network to obtain a second current sub-image feature, wherein the first image feature includes the first current sub-image feature, and the second image feature includes the second current sub-image feature.

9. According to any one of the methods of claim 1 to 8, before acquiring the image to be analyzed obtained by performing image acquisition on the part object to be inspected, the method further comprises: Acquire K positive sample images including the normal part object and K negative sample images including the abnormal part object to obtain K sample image pairs, wherein the negative sample images are synthesized after adjusting the positive sample images, and K is a natural number greater than 1; The K sample images are used to train the initialized abnormal part recognition network until a comparison result between the output features of the second feature extraction network and the output features of the first feature extraction network reaches a convergence condition.

10. According to the method of claim 9, the training of the initialized abnormal part recognition network using the K sample images until the comparison result between the output features of the second feature extraction network and the output features of the first feature extraction network reaches a convergence condition comprises: Obtain a current sample image pair from the K sample image pairs, and perform the following operations: input a current positive sample image in the current sample image pair into the first feature extraction network to obtain a current positive sample sub-image feature, and input a current negative sample image in the current sample image pair into the second feature extraction network to obtain a current negative sample sub-image feature; When the feature similarity between the current negative sample sub-image feature and the current positive sample sub-image feature is less than a third threshold, obtaining a next pair of sample images as the current sample image pair; When the feature similarity between the feature of the current negative sample sub-image and the feature of the current positive sample sub-image is greater than or equal to a third threshold, adding one to the training convergence statistics result; When the training convergence statistical result reaches a fourth threshold, it is determined that the convergence condition is met.

11. According to the method of claim 9 or 10, the step of acquiring K positive sample images including the normal part objects and K negative sample images including the abnormal part objects to obtain K sample image pairs comprises: Acquire K positive sample images including the normal part object; The K positive sample images are adjusted to synthesize the K negative sample images, wherein the positive sample images are sample template images including the normal part objects.

12. According to the method of claim 11, the step of adjusting the K positive sample images to synthesize the K negative sample images comprises one of the following: Adding white noise to the positive sample image to generate the negative sample image; Get multiple abnormal part image blocks, where: The abnormal part image block includes part defects that belong to the same type of parts as the part object to be detected; one or at least two abnormal part image blocks among the multiple abnormal part image blocks are overlaid on the positive sample image to generate the negative sample image; The positive sample image is input into an abnormal part generation network to obtain the negative sample image, wherein the abnormal part generation network is a network obtained by training with the positive sample image and abnormal part image blocks and used to generate an image including the abnormal part object, and the abnormal part image block includes part defects belonging to the same type of parts as the part object to be detected.

13. An image recognition device, the device being deployed on an electronic device, comprising: A first acquisition unit is used to acquire an image to be analyzed obtained by collecting an image of the part object to be inspected; A recognition unit, used for extracting image features of the image to be analyzed by using an abnormal part recognition network having a twin network structure, wherein the image features include a first image feature extracted from the image to be analyzed by a first feature extraction network in the twin network structure, and a second image feature extracted from the image to be analyzed by a second feature extraction network in the twin network structure, the first feature extraction network is trained using a positive sample image, the second feature extraction network is trained using a negative sample image with the first feature extraction network as a reference network, the positive sample image is a sample image including a normal part object in a sample image pair, and the negative sample image is a sample image including an abnormal part object synthesized based on the positive sample image; A second acquisition unit, configured to acquire a feature similarity between the first image feature and the second image feature; A determination unit is used to determine a recognition result of the image to be analyzed based on the feature similarity, wherein the recognition result is used to indicate whether the part object to be detected in the image to be analyzed is the abnormal part object.

14. A computer-readable storage medium, the computer-readable storage medium comprising a stored program, wherein: When the program is executed by a processor, the method described in any one of claims 1 to 12 is executed.

15. A computer program product, comprising a computer program, which implements the steps of the method according to any one of claims 1 to 12 when executed by a processor.

16. An electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the method according to any one of claims 1 to 12 through the computer program.