Defect detection method, device, apparatus, and storage medium device
Patent Information
- Application Number
- CN202310092571.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-16
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2043-01-16
AI Technical Summary
[0019] On the other hand, embodiments of this application provide a computer program product, which includes a computer program that, when executed by a processor, implements the defect detection method provided in embodiments of this application.
Smart Images

Figure CN116977255B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and in particular to a defect detection method, apparatus, device, and storage medium. Background Technology
[0002] As living standards continue to improve, all industries need to conduct defect detection on products to reduce defects and thus meet people's increasingly normal material needs.
[0003] Current defect detection methods primarily rely on manual visual inspection to identify defects, or on comparing images of the product with standard images to determine defects. While existing technologies can detect defects in products to some extent, their accuracy is low and their efficiency is limited. Therefore, how to efficiently and accurately determine whether a product contains defects has become an urgent problem to be solved. Summary of the Invention
[0004] This application provides a defect detection method, apparatus, device, and storage medium, which can efficiently and accurately determine whether the object to be detected has defects and has high applicability.
[0005] On one hand, embodiments of this application provide a defect detection method, the method comprising:
[0006] An initial image of the object to be detected and a standard image of a standard object are determined, wherein the initial image and the standard image are obtained based on the same shooting conditions;
[0007] Determine the initial image features of the initial image and the standard image features of the standard image.
[0008] Determine the offset feature value of each pixel of the initial image feature relative to the pixel at the same position in the standard image feature, and determine the first fused image feature based on the initial image feature and the offset feature value corresponding to each pixel of the initial image feature;
[0009] Based on the first fused image features, the object type of the object to be detected is determined. The object type includes a first type and a second type. When the object to be detected belongs to the first type, the object to be detected has defects compared to the standard object. When the object to be detected belongs to the first type, the object to be detected does not have defects compared to the standard object.
[0010] On the other hand, embodiments of this application provide a defect detection device, which includes:
[0011] The image determination module is used to determine the initial image of the object to be detected and the standard image of the standard object. The initial image and the standard image are obtained based on the same shooting conditions.
[0012] The feature extraction module is used to determine the initial image features of the initial image and the standard image features of the standard image.
[0013] The feature processing module is used to determine the offset feature value of each pixel of the initial image feature relative to the pixel at the same position in the standard image feature, and to determine the first fused image feature based on the initial image feature and the offset feature value corresponding to each pixel of the initial image feature.
[0014] The type determination module is used to determine the object type of the object to be detected based on the first fused image features. The object type includes a first type and a second type. When the object to be detected belongs to the first type, the object to be detected has defects compared to the standard object. When the object to be detected belongs to the first type, the object to be detected does not have defects compared to the standard object.
[0015] On the other hand, embodiments of this application provide an electronic device, including a processor and a memory, which are interconnected;
[0016] The aforementioned memory is used to store computer programs;
[0017] The processor described above is configured to execute the defect detection method provided in the embodiments of this application when the computer program described above is invoked.
[0018] On the other hand, embodiments of this application provide a computer-readable storage medium storing a computer program that is executed by a processor to implement the defect detection method provided in embodiments of this application.
[0019] On the other hand, embodiments of this application provide a computer program product, which includes a computer program that, when executed by a processor, implements the defect detection method provided in embodiments of this application.
[0020] In this embodiment, after determining the initial image features of the initial image of the object to be detected and the standard image features of the standard image of the standard object, the first fused image features can be determined by using the offset feature value of each pixel of the initial image features compared to the pixel at the same position in the standard image features, and the initial image features. This includes the differences between the object to be detected and the standard object caused by its own deformation, as well as the differences between the defects of the object to be detected and the standard object. Thus, considering the differences caused by the deformation of the object to be detected, the presence of defects in the image to be detected can be determined efficiently and accurately based on the first fused image features, which has high applicability. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a schematic diagram of the network structure of the defect detection method provided in the embodiments of this application;
[0023] Figure 2 This is a flowchart illustrating the defect detection method provided in the embodiments of this application;
[0024] Figure 3 This is a schematic diagram of a scene for determining image features provided in an embodiment of this application;
[0025] Figure 4 This is a schematic diagram of the structure of the residual unit provided in the embodiment of this application;
[0026] Figure 5 This is a schematic diagram of a scenario for determining the first offset feature provided in an embodiment of this application;
[0027] Figure 6 This is a schematic diagram of a scenario for determining the second offset feature provided in an embodiment of this application;
[0028] Figure 7 This is a schematic diagram of a scene for determining the features of the first fused image, provided in an embodiment of this application;
[0029] Figure 8 This is a schematic diagram of a scenario for determining connected components provided in an embodiment of this application;
[0030] Figures 9a-9d This is a schematic diagram of a defect detection scenario provided in an embodiment of this application;
[0031] Figure 10a This is a schematic diagram of a scenario for fabric defect detection provided in an embodiment of this application;
[0032] Figure 10b This is another schematic diagram of a fabric defect detection scenario provided in the embodiments of this application;
[0033] Figure 11 This is a flowchart illustrating the training method of the defect detection model provided in the embodiments of this application;
[0034] Figure 12 This is a schematic diagram of the structure of the initial model provided in the embodiments of this application;
[0035] Figure 13This is a schematic diagram of the defect detection device provided in the embodiments of this application;
[0036] Figure 14 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0037] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0038] The defect detection method provided in this application can detect defects in objects in any field. For example, in the industrial production field, the defect detection method provided in this application can determine whether products produced on an assembly line have defects, such as whether textile fabrics have pattern errors, cross-cut defects, bad ground, spots, or damage. As another example, in fields such as mapping and transportation, in Intelligent Traffic Systems (ITS) or Intelligent Vehicle Infrastructure Cooperative Systems (IVICS), the defect detection method provided in this application can determine whether traffic elements such as traffic lights and road surfaces are damaged, thereby determining the road traffic status of each road in the road network at each time, thus providing technical support for accurate navigation.
[0039] Intelligent Transportation Systems (ITS), also known as Intelligent Transportation Systems, effectively integrate advanced science and technology (information technology, computer technology, data communication technology, sensor technology, electronic control technology, automatic control theory, operations research, artificial intelligence, etc.) into transportation, service control, and vehicle manufacturing. This strengthens the connection between vehicles, roads, and users, thereby forming a comprehensive transportation system that ensures safety, improves efficiency, enhances the environment, and saves energy.
[0040] Among them, the intelligent vehicle-road cooperative system, or vehicle-road cooperative system for short, is a development direction of intelligent transportation systems (ITS). The vehicle-road cooperative system uses advanced wireless communication and next-generation Internet technologies to implement dynamic real-time information interaction between vehicles and roads in all aspects. Based on the collection and fusion of dynamic traffic information in all time and space, it carries out active safety control of vehicles and cooperative road management, fully realizing effective coordination between people, vehicles, and roads, ensuring traffic safety, improving traffic efficiency, and thus forming a safe, efficient, and environmentally friendly road traffic system.
[0041] See Figure 1 , Figure 1 This is a schematic diagram of the network structure of the defect detection method provided in the embodiments of this application. For example... Figure 1 As shown, for any initial image 11 of an object to be inspected, a standard image 12 of a corresponding standard object can be determined. The standard object corresponding to the object to be inspected is a reference object without defects.
[0042] After determining the initial image 11 of the object to be detected and the standard image 12 of the standard object, the object type of the object to be detected can be determined based on the terminal 13.
[0043] The object type includes a first type and a second type. When the object type of the object to be tested belongs to the first type, it means that the object to be tested has a defect compared to the standard object. When the object type of the object to be tested belongs to the second type, it means that the object to be tested does not have a defect compared to the standard object.
[0044] In this embodiment, device 13 can be a terminal device or a server, which can be determined based on the actual application scenario requirements and is not limited here.
[0045] The server can be a vehicle network server or other independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms, such as vehicle network edge cloud platforms and cloud computing platforms.
[0046] The terminal device can be a smartphone, tablet, laptop, desktop computer, smart voice interaction device (such as a smart speaker), wearable electronic device (such as a smartwatch), in-vehicle terminal, smart home appliance (such as a smart TV), AR / VR device, etc., but is not limited to these.
[0047] See Figure 2 , Figure 2 This is a schematic flowchart of the defect detection method provided in the embodiments of this application. Figure 2 As shown, the defect detection method provided in this application embodiment may specifically include the following steps:
[0048] Step S21: Determine the initial image of the object to be detected and the standard image of the standard object.
[0049] In this embodiment, the standard object is a reference object without defects, and the object to be inspected is an object that is the same as the standard object and requires defect detection. For example, in the textile fabric production process, the newly produced textile fabric from the production line is the object to be inspected, and the corresponding sample fabric is the standard object.
[0050] In this embodiment of the application, the initial image of the object to be detected and the standard image of the standard object are obtained based on the same shooting conditions.
[0051] The aforementioned shooting conditions include, but are not limited to, one or more of the following: shooting angle, shooting equipment, equipment parameters of the shooting equipment, shooting background, shooting distance, or lighting conditions. The specific conditions can be determined based on the actual application scenario requirements and are not limited here.
[0052] Step S22: Determine the initial image features of the initial image and the standard image features of the standard image.
[0053] In some feasible implementations, when determining the initial image features of the initial image and the standard image features of the standard image, the initial image and the standard image can be processed based on a pre-trained feature extraction network to obtain the initial image features of the initial image and the standard image features of the standard image.
[0054] Optionally, since the object to be detected may undergo certain deformation due to external force factors or its own material when determining the initial image of the object to be detected, the aforementioned deformation will have a significant impact on the extraction of the initial image features when determining the initial image features based on the same feature extraction network. This will result in a significant difference between the initial image features obtained under the condition of deformation and the initial image features obtained under the condition of no deformation, thereby affecting the defect detection effect on the object to be detected.
[0055] Therefore, when determining the initial image features of the initial image, the initial image features can be obtained by processing the initial image using a first feature extraction network. When determining the standard image features of the standard image, the standard image features can be obtained by processing the standard image using a second feature extraction network.
[0056] The first feature extraction network and the second feature extraction network have the same network structure.
[0057] The network weights of the first feature extraction network are determined based on the network weights of the second feature extraction network and preset weight components.
[0058] The network weights of the second feature extraction network are the network weights that are finally determined during the network training process. The preset weights can also be the weight components determined during the simultaneous training of the first and second feature extraction networks, or they can be pre-set weight components. The specific weights can be determined based on the actual application scenario requirements, and there are no restrictions here.
[0059] As an example, the network weights of the first feature extraction network and the network weights of the second feature extraction network have the following relationship:
[0060] ω1=ω2+Δω
[0061] Where ω1 is the network weight of the first feature extraction network, ω2 is the network weight of the second feature extraction network, and Δω is a preset weight component.
[0062] The network weights of each network layer or structure in the first feature extraction network and the corresponding network weights of the network layer or structure in the second feature extraction network have the above-mentioned relationship.
[0063] See Figure 3 , Figure 3 This is a schematic diagram of a scene for determining image features provided in an embodiment of this application. The initial image of the object to be detected and the standard image of the standard object are shown below. Figure 3 In the case shown, the initial image of the object to be detected can be input into the first feature extraction network to obtain the initial image features of the object to be detected, and the standard image of the standard object can be input into the second feature extraction network to obtain the standard image features of the standard object.
[0064] Wherein, when the network weight of the first feature extraction network is ω1 and the network weight of the second feature extraction network is ω2, the relationship between the network weight of the first feature extraction network and the network weight of the second feature extraction network is ω1=ω2+Δω, where Δω is a preset weight component.
[0065] Given the aforementioned relationship between the network weights of the first and second feature extraction networks, they can extract image features in a similar manner while maintaining their respective learning degrees of freedom to achieve better feature representation capabilities. Furthermore, they can mitigate the impact of potential differences between non-defect regions in the initial image and the standard image on defect detection, thereby improving defect detection accuracy.
[0066] In the embodiments of this application, the feature extraction network used to process the features of the initial image and the standard image includes, but is not limited to, Residual Network (ResNet), Visual Geometry Group (VGG) network, etc. The specific network can be determined based on the actual application scenario requirements and is not limited here.
[0067] The residual networks include, but are not limited to, ResNet18, ResNet34, ResNet50, and ResNet101. The specific network can be determined based on the actual application scenario requirements, and no restrictions are imposed here.
[0068] Taking the ResNet50 network as an example, ResNet50 consists of multiple concatenated residual blocks. Concatenation means that the output of the previous residual block is used as the input of the current residual block. Each residual block contains multiple convolutional layers, with adjacent concatenated layers and skip connections between non-adjacent layers, thus forming a residual block with a residual structure. Each residual block is composed of consecutive residual units. For each residual unit, channel downsampling is first performed, followed by feature transformation using a 3x3 convolution, then channel upsampling to restore the original channel size, and finally, output features are obtained through residual connections with the input. Each residual unit achieves spatial downsampling through max pooling layers or convolutional layers with a stride of 2, increasing the network's receptive field while maintaining local translation invariance. Simultaneously, the baseline number of channels for each residual unit increases with network depth, allowing for the extraction of richer semantic information.
[0069] For example, a residual block in a ResNet50 network contains six convolutional layers. Besides these six layers being interconnected, there are also skip connections from the output of the previous residual block (i.e., the input of the first convolutional layer) to the input of the fourth convolutional layer. That is, after the output of the previous residual block is fused with the output of the third convolutional layer, the resulting feature map is input to the fourth convolutional layer. Furthermore, there are skip connections from the input of the fourth convolutional layer to the output of the sixth convolutional layer. That is, after the input of the fourth convolutional layer (referring to the feature map obtained by fusing the output of the third convolutional layer with the output of the previous residual block) is fused with the output of the sixth convolutional layer, the resulting feature map is used as the output of the current residual block and input to the next residual block. The internal structure of other residual blocks in the ResNet50 network is similar to the example above and will not be elaborated upon here.
[0070] See Figure 4 , Figure 4 This is a schematic diagram of the structure of the residual unit provided in an embodiment of this application. For example... Figure 4 As shown, assuming x is the input feature map, the output of the residual unit is H(x) = F(x) + x. Here, F(x) is the residual, the weight layer in the figure represents the convolution operation, and generally a residual part contains 2-3 convolution operations. The convolutional feature map F(x) is added to x to obtain a new feature map H(x), and ReLU represents the activation function.
[0071] Step S23: Determine the offset feature value of each pixel of the initial image feature relative to the pixel at the same position in the standard image feature, and determine the first fused image feature based on the initial image feature and the offset feature value corresponding to each pixel of the initial image feature.
[0072] In some feasible implementations, for the initial image features corresponding to the object to be detected and the standard image features corresponding to the standard object, the differences between the two feature image features may include the differences between the object to be detected and the standard object caused by the deformation of the object itself, as well as the differences between the defects existing in the object to be detected and the standard object.
[0073] For example, the difference between the standard image features of the standard image and the initial image features of the initial image of the initial textile fabric produced with reference to the standard textile fabric can include the feature differences caused by the fabric's unevenness, wrinkles, etc. due to the flexible deformation of the textile fabric when taking the images corresponding to the standard textile fabric and the initial textile fabric, and can also include the feature differences caused by the inherent defects of the initial textile fabric (such as holes).
[0074] Therefore, when performing defect detection on the object to be tested, it is necessary to further consider the differences between the object to be tested and the standard object caused by its own deformation.
[0075] Specifically, the offset feature value of each pixel in the initial image feature relative to the pixel at the same position in the standard image feature can be determined. This offset feature value characterizes the feature offset between each pixel in the initial image feature and the pixel at the same position in the standard image feature when the object to be detected is deformed.
[0076] Furthermore, the first fused image features can be determined based on the initial image features and the offset feature values corresponding to each pixel of the initial image features. The first fused image features are the feature information corresponding to the object to be detected after considering the deformation of the object to be detected, that is, the feature information corresponding to the object to be detected after eliminating the feature differences caused by the deformation of the object to be detected itself.
[0077] For example, for a standard image of a standard textile fabric without defects, and an initial image of an initial textile fabric produced with reference to the standard textile fabric, the offset feature value of each pixel of the initial image feature of the initial textile fabric relative to the pixel at the same position in the standard image feature of the standard textile fabric can be determined. This offset feature value characterizes the feature offset between each pixel of the initial image feature and the pixel at the same position in the corresponding standard image feature when the initial textile fabric undergoes flexible deformation.
[0078] Based on this, the first fused image feature is determined based on the initial image features of the initial textile fabric and the offset feature value corresponding to each pixel of the initial image features. At this time, the first fused image feature is the feature information obtained under the condition that the initial textile fabric undergoes flexible deformation, that is, the feature information of the initial textile fabric after eliminating the difference between the initial textile fabric itself and the standard image features caused by the flexible deformation of the initial textile fabric.
[0079] In some feasible implementations, when determining the offset feature value of each pixel of the initial image feature relative to the pixel at the same position in the standard image feature, the offset position of each pixel of the initial image feature relative to the pixel at the same position in the standard image feature can be determined first based on the initial image feature and the standard image feature. Then, the target feature value corresponding to the offset position of each pixel of the initial image feature in the standard image feature can be determined, and the target feature value corresponding to each pixel of the initial image feature can be determined as the offset feature value relative to the pixel at the same position in the standard image feature.
[0080] Specifically, the second fused image features can be determined based on the initial image features and the standard image features. For example, the initial image features and the standard image features can be concatenated by channels, and the concatenated features can be used as the second fused image features. Furthermore, the second fused image features can be processed based on an offset prediction network to predict the offset position of each pixel in the initial image features relative to the pixel at the same position in the standard image features.
[0081] Optionally, a second fused image feature can be determined first based on the initial image feature and the standard image feature, and then the offset information of each pixel of the initial image feature relative to the pixel at the same position in the standard image feature can be determined based on the second fused image feature.
[0082] For each pixel in the initial image features, the offset information for that pixel includes the horizontal offset amount, the horizontal offset direction, the vertical offset direction, and the vertical offset amount. The horizontal offset direction indicates whether the pixel shifts to the left or right, the vertical offset direction indicates whether the pixel shifts upward or downward, the horizontal offset amount indicates the distance the pixel shifts to the left or right, and the vertical offset amount indicates the distance the pixel shifts upward or downward.
[0083] Furthermore, for each pixel of the initial image feature, after determining the offset information of the pixel, the offset position of the pixel relative to the pixel at the same position in the standard image feature can be determined based on the offset information of the pixel.
[0084] In some feasible implementations, when determining the offset information of each pixel of the initial image feature relative to the pixel at the same position in the standard image feature based on the second fused image feature, a first offset feature can be first determined based on the second fused image feature. Each pixel of the first offset feature is used to characterize the lateral offset reference value of the pixel at the same position in the initial image feature. For each pixel of the initial image feature, the lateral offset amount and lateral offset direction of the pixel relative to the pixel at the same position in the standard image feature can be determined based on the lateral offset reference value of that pixel.
[0085] See Figure 5 , Figure 5 This is a schematic diagram of a scenario for determining the first offset feature provided in an embodiment of this application. For example... Figure 5 As shown, when determining the first offset feature, the second fused image features can be convolved based on the first convolutional network to obtain the first initial offset feature, and then the feature value of each pixel of the first initial offset feature can be normalized to obtain the first offset feature.
[0086] The first convolutional network includes at least one cascaded convolutional layer, and the configuration information of each convolutional layer can be determined based on the actual application scenario requirements, without any restrictions.
[0087] Based on this, the feature value of each pixel of the first offset feature is the horizontal offset reference value of the pixel at the same position in the initial image feature, and the feature value of each pixel of the first offset feature is greater than or equal to 0 and less than or equal to 1.
[0088] Similarly, a second offset feature can be determined based on the second fused image features. Each pixel of the second offset feature is used to characterize the vertical offset reference value of a pixel at the same position in the initial image features. For each pixel of the initial image features, the vertical offset amount and direction of the pixel relative to a pixel at the same position in the standard image features can be determined based on the vertical offset reference value of that pixel.
[0089] See Figure 6 , Figure 6 This is a schematic diagram of a scenario for determining the second offset feature provided in an embodiment of this application. For example... Figure 6 As shown, when determining the second offset feature, the second initial offset feature can be obtained by performing convolution processing on the second fused image feature based on the second convolution network, and then the feature value of each pixel of the second initial offset feature can be normalized to obtain the second offset feature.
[0090] The second convolutional network includes at least one cascaded convolutional layer, and the parameter information of each convolutional layer can be determined based on the actual application scenario requirements, without any restrictions.
[0091] The network parameters of the first and second convolutional networks are not the same, but this is not a restriction.
[0092] Based on this, the feature value of each pixel in the second offset feature is the vertical offset reference value of the pixel at the same position in the initial image feature, and the feature value of each pixel in the second offset feature is greater than or equal to 0 and less than or equal to 1.
[0093] For each pixel in the initial image features, when determining the lateral offset and direction of that pixel relative to pixels at the same position in the standard image features based on its lateral offset reference value, the lateral offset reference value can be compared with a first threshold. If the lateral offset reference value is less than the first threshold, the lateral offset direction of that pixel is determined to be to the left; if the lateral offset reference value is greater than the first threshold, the lateral offset direction of that pixel is determined to be to the right; if the lateral offset reference value is greater than the first threshold, then that pixel does not shift laterally.
[0094] The first threshold can be determined based on the actual application scenario requirements, such as 0.5, and is not limited here.
[0095] For each pixel in the initial image features, when determining the lateral offset of that pixel, the absolute value of the difference between the lateral offset reference value and the first threshold can be determined. For ease of description, the absolute value of the difference between the lateral offset reference value and the first threshold is referred to as the first absolute value. Based on this, the lateral offset of the pixel can be determined based on the first absolute value and the lateral offset coefficient, such as by multiplying the first absolute value and the lateral offset coefficient to determine the lateral offset of the pixel.
[0096] The lateral offset coefficient is used to characterize the horizontal mapping relationship between each pixel of the first offset feature and the receptive field of the standard image.
[0097] Based on the above implementation method, it can be determined whether each pixel of the initial image feature needs to be moved to the left or right, and the lateral offset of moving to the left or to the right.
[0098] For each pixel in the initial image features, when determining the vertical offset and direction of that pixel relative to pixels at the same position in the standard image features based on its vertical offset reference value, the vertical offset reference value can be compared with a second threshold. If the vertical offset reference value is less than the second threshold, the vertical offset direction of that pixel is determined to be upward; if the vertical offset reference value is greater than the second threshold, the vertical offset direction of that pixel is determined to be downward; if the vertical offset reference value is greater than the second threshold, then that pixel is not offset vertically.
[0099] The second threshold can be determined based on the actual application scenario requirements, such as 0.5, and is not limited here.
[0100] For each pixel in the initial image features, when determining the vertical offset of that pixel, the absolute value of the difference between the vertical offset reference value and the second threshold can be determined. For ease of description, the absolute value of the difference between the vertical offset reference value and the second threshold is referred to as the second absolute value. Based on this, the vertical offset of the pixel can be determined based on the second absolute value and the vertical offset coefficient, such as determining the vertical offset of the pixel by multiplying the second absolute value and the vertical offset coefficient.
[0101] The longitudinal offset coefficient is used to characterize the longitudinal mapping relationship between each pixel of the second offset feature and the receptive field of the standard image.
[0102] Based on the above implementation method, it can be determined whether each pixel of the initial image feature needs to be moved up or down, and the vertical offset of the upward or downward movement.
[0103] Specifically, for each pixel in the initial image features, when determining the feature value of the pixel relative to the pixel at the same position in the standard image features, a bilinear interpolation can be performed on the offset position of the pixel in the standard image features based on the standard image features to obtain the target feature value of the offset position of the pixel in the standard image features, and this target feature value is determined as the offset feature value of the pixel relative to the pixel at the same position in the standard image features.
[0104] In some feasible implementations, after determining the offset feature value corresponding to each pixel of the initial image feature, the feature value corresponding to each pixel of the initial image feature and the corresponding offset feature value can be fused to obtain a third fused image feature.
[0105] Specifically, the third fused image feature can be obtained by concatenating each pixel of the initial image feature with its corresponding offset feature value, or by performing convolution processing on the initial image feature and the feature map obtained with each offset feature value based on a convolutional network, and concatenating or adding the feature values of pixels at the same position to obtain the third fused image feature.
[0106] Alternatively, for each pixel in the initial image features, the feature value of the pixel and the corresponding offset feature value can be added to obtain the actual sampling deviation of the pixel. Then, a convolution operation is performed on the initial image features, and the actual sampling deviation of each pixel is used as the convolution bias to finally obtain the third fused image features.
[0107] The third fused image feature is the feature obtained by aligning the features of the standard object and the object to be detected at the same object position.
[0108] Furthermore, after determining the third fused image feature, the first fused image feature can be determined based on the third fused image feature and the initial image feature. For example, the difference between the feature values of each pixel of the third fused image feature and the initial image feature can be determined as the feature value of each pixel of the first fused image feature to obtain the first fused image feature. Alternatively, after convolving the third fused image feature based on at least one cascaded convolutional network, the difference between the feature value of each pixel of the convolved feature map and the feature value of the corresponding pixel in the initial image feature can be determined as the feature value of each pixel of the first fused image feature to obtain the first fused image feature.
[0109] The following is combined Figure 7 The process for determining the first fused image features is further explained. Figure 7This is a schematic diagram of a scenario for determining the first fused image features provided in an embodiment of this application. The standard image shown in Figure X is an image of a standard textile fabric (standard object), and the initial image is an image of a newly produced initial textile fabric (object to be detected) with reference to the standard image.
[0110] The initial image is input into the first feature extraction network to obtain the initial image features of the initial textile fabric, and the standard image is input into the second feature extraction network to obtain the standard image features of the standard textile fabric. The network weights of the first and second feature extraction networks are related as ω1 = ω2 + Δω, where Δω is a preset weight component, the network weight of the first feature extraction network is ω1, and the network weight of the second feature extraction network is ω2.
[0111] The flexible alignment module can determine the third fused image features by flexibly aligning the features of the initial textile fabric and the standard textile fabric at the same fabric position, based on the initial image features and the standard image features.
[0112] Furthermore, based on the third fused image features and the initial image features, the first fused image features can be determined. The first fused image features are the feature information of the initial textile fabric after eliminating the differences between the initial textile fabric and the standard image caused by the flexible deformation of the initial textile fabric itself.
[0113] The implementation method for determining the third fusion feature by the flexible alignment module can be found in the implementation methods shown in the aforementioned embodiments, and will not be repeated here.
[0114] Step S24: Determine the object type of the object to be detected based on the features of the first fused image.
[0115] The object types in this application embodiment include a first type and a second type. When any object belongs to the first type, it means that the object has a defect compared to the corresponding standard object. When any object belongs to the second type, it means that the object does not have a defect compared to the corresponding standard object, that is, it is consistent with the corresponding standard object.
[0116] In some feasible implementations, after determining the first fused image features based on the initial image features of the object to be detected and the standard image features of the standard object, the type features corresponding to the object to be detected can be determined based on the first fused image features.
[0117] Specifically, the first fused image features can be pooled based on the pooling layer, and the pooled features can be further processed based on the fully connected layer to obtain the type features corresponding to the object to be detected.
[0118] Among them, the pooling processing of the first fused image features based on the pooling layer includes, but is not limited to, global pooling, average pooling, etc. The specific method can be determined based on the actual application scenario requirements and is not limited here.
[0119] Furthermore, based on the type characteristics of the object to be detected, the type probability of the object can be determined. When the type probability of the object to be detected is greater than or equal to the probability threshold, the object type of the object to be detected is determined to be the first type, that is, the object to be detected has a defect compared to the standard object. When the type probability of the object to be detected is less than the probability threshold, the object type of the object to be detected can be determined to be the second type, that is, the object to be detected has no defect compared to the standard object, that is, the object to be detected is consistent with the standard object.
[0120] It should be emphasized that any defect in the object to be tested in the embodiments of this application indicates that the object to be tested has a difference from the standard object. In other words, the defect in the object to be tested is the difference from the standard object, such as product stains, gaps, damage, or fabric defects, etc., which are not limited here.
[0121] In some feasible implementations, after determining the object type to which the object to be detected belongs, if the object to be detected belongs to the first type, that is, if it is determined that the object to be detected has defects compared with the standard object, the target defects in the image to be detected can be determined based on the initial image of the object to be detected, and the target defects can be classified.
[0122] Specifically, a heatmap of the initial image of the object to be detected can be determined, and the heatmap can be binarized to obtain the target grayscale image.
[0123] The heatmap of the initial image of the object to be detected can be determined based on the Class Activation Mapping (CAM) model.
[0124] Image binarization involves setting the grayscale value of points in an image to either 0 or 255, thus presenting the entire image in a distinct black and white manner. Essentially, it involves obtaining a binarized image from a grayscale image with 256 brightness levels by applying an appropriate threshold, while still retaining the overall and local features of the image.
[0125] Furthermore, such as Figure 8 As shown, Figure 8 This is a schematic diagram illustrating a scenario for determining connected components, as provided in an embodiment of this application. For example... Figure 8As shown, after determining the target grayscale image, each connected component in the target grayscale image can be determined, that is, the image region composed of pixels with the same pixel value and adjacent positions in the target grayscale image can be determined, and each connected component is defined as a candidate defect region. Each candidate defect region is the region in the initial image of the object to be detected that may have defects.
[0126] Based on this, the target defect in the object to be detected can be determined from the image regions corresponding to each connected component in the initial image of the object to be detected. For each connected component, the image region corresponding to that connected component in the initial image of the object to be detected can be compared with the image region corresponding to that connected component in the standard image. If the two are the same or the image similarity is greater than a similarity threshold, it means that the image region corresponding to that connected component in the initial image of the object to be detected does not contain the target defect. If the two are the same or the image similarity is less than or equal to the similarity threshold, it means that the image region corresponding to that connected component in the initial image of the object to be detected contains the target defect.
[0127] In this process, after determining the heatmap of the initial image and converting it to grayscale using binarization or other grayscale conversion methods, regions in the initial image that may contain defects can be identified based on the connected components in the grayscale image. Figures 9a-9d For any grayscale image shown, multiple connected components can be determined based on the grayscale image, and then the image region that ultimately includes the defect in the initial image corresponding to the grayscale image can be determined in each connected component.
[0128] For any initial image of an object to be inspected, the area containing defects can be determined based on the heat map. This allows for visualization of the defect area once the defect is identified, enabling the visualization and effective localization of the defect.
[0129] In this embodiment, the initial image features of the initial image of the object to be detected and the standard image features of the standard image of the standard object are determined by the first feature extraction network and the second feature extraction network, respectively. This allows for accurate extraction of image features from each image while adapting to feature points such as deformations that occur within the object itself. By using the offset feature value of each pixel in the initial image features compared to the pixel at the same position in the standard image features, and the initial image features, a first fused image feature is determined, including the differences between the object to be detected and the standard object caused by its own deformation, as well as the differences between defects in the object to be detected and the standard object. This achieves feature alignment between the object to be detected and the standard object. Therefore, based on the first fused image feature, the presence of defects in the image to be detected can be efficiently and accurately determined while eliminating differences caused by the deformation of the object itself, demonstrating high applicability.
[0130] For example, in the case of textile fabrics, if the initial image of the textile fabric to be tested is taken when the textile fabric to be tested has undergone flexible deformation, the defect detection method provided by the embodiments of this application can eliminate the difference between the fabric to be tested and the standard fabric caused by the flexible deformation, so that the presence of defects in the fabric to be tested can be accurately and efficiently determined regardless of whether the fabric to be tested has undergone flexible deformation.
[0131] like Figure 10a As shown, Figure 10a This is a schematic diagram of fabric defect detection provided in an embodiment of this application. Figure 10a When the fabric to be tested has worn yarns, regardless of whether the fabric undergoes flexible deformation when the image of the fabric to be tested is captured, the defect detection method provided in this application embodiment can accurately detect the relevant areas of worn yarns in the fabric to be tested compared to the standard fabric, and effectively detect defects in the fabric to be tested.
[0132] like Figure 10b As shown, Figure 10b This is another schematic diagram of fabric defect detection provided in an embodiment of this application. Figure 10b When the fabric to be tested has holes (defects that are not dyed or damaged), regardless of whether the fabric undergoes flexible deformation when the image of the fabric is taken, the defect detection method provided in this application can detect small, hard-to-observe holes in the fabric compared to the standard fabric.
[0133] In some feasible implementations, the defect detection method provided in this application embodiment can be implemented by a defect detection model, that is... Figure 2 Any defect detection method can be implemented using a defect detection model. The defect detection model is based on... Figure 11 The training method shown was used to train the equipment.
[0134] See Figure 11 , Figure 11 This is a flowchart illustrating the training method of the defect detection model provided in this application embodiment. For example... Figure 11 As shown, the training method for the defect detection model provided in this application embodiment may specifically include the following steps:
[0135] Step S111: Determine the training sample set. The training sample set includes multiple sets of training samples. Each set of training samples includes an initial image of a sample object and a standard image of the standard object corresponding to that sample object.
[0136] In some feasible implementations, when determining the training sample set, a standard image of at least one standard object can be determined. Further, for each standard object, different defects can be added to the standard object to obtain at least one defective sample object corresponding to that standard object, while simultaneously determining an initial image for each sample object.
[0137] Based on this, for each standard object, the standard image of the standard object and the initial image of a corresponding sample object can be determined as a set of training samples, or two standard images of the standard object can be determined as a set of training samples, with one standard image serving as the initial image of a sample object.
[0138] Each standard object can be any object such as textile fabric, leather, poster, painting, and paper, and the specific object can be determined based on the actual application scenario requirements, without any restrictions here.
[0139] The aforementioned preset storage space can be a remote dictionary service (Redis) system, a database, cloud storage, or a blockchain. The specific location can be determined based on the actual application scenario requirements and is not limited here.
[0140] After the training sample set is determined, each group of training samples in the training sample set can be stored in a designated storage space. This designated storage space includes, but is not limited to, cloud storage, databases, blockchain, and the storage space of the device itself that performs the model training task. The specific storage space can be determined based on the actual application scenario requirements and is not limited here.
[0141] In short, a database can be viewed as an electronic filing cabinet—a place to store electronic files. It can be a relational database (SQL database) or a non-relational database (NoSQL database), without limitation. In this application, it can be used to store various training samples in the training sample set. Blockchain is a new application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked using cryptographic methods. In the embodiments of this application, each data block in the blockchain can store various training samples in the training sample set. Cloud storage is a new concept extended and developed from cloud computing. It refers to the use of cluster applications, grid technology, and distributed storage file systems to aggregate a large number of various types of storage devices (also called storage nodes) in a network through application software or application interfaces to work together and jointly store various training samples in the training sample set.
[0142] Step S112: Input each group of training samples into the initial model to predict the object type of the sample object in each group of training samples based on the initial model.
[0143] In some feasible implementations, for each set of training samples, the predicted object type of the sample object in the training sample also includes a first type and a second type. When the sample object is predicted to belong to the first type, it means that the sample object has a defect compared to the corresponding standard object. When the sample object is predicted to belong to the second type, it means that the sample object does not have a defect compared to the corresponding standard object.
[0144] When training the initial model based on the training sample set, each set of training samples can be input into the initial model. The initial model processes the initial image and standard image in each set of training samples to predict the object type of the sample object in each set of training samples.
[0145] The initial model may include a feature processing network, a feature alignment network, and a type prediction network, see [link to relevant documentation]. Figure 12 , Figure 12 This is a schematic diagram of the structure of the initial model provided in the embodiments of this application.
[0146] like Figure 12 As shown, the feature processing network may include a first feature extraction network and a second feature extraction network. For each training sample, the feature processing network determines the initial image features of the sample object in the training sample through the first feature extraction network and determines the standard image features of the standard object corresponding to the training sample through the second feature extraction network. The relationship between the network weights ω1 of the first feature extraction network and the network weights ω2 of the second feature extraction network is ω1 = ω2 + Δω, where Δω is a preset weight component.
[0147] For details on the implementation of the feature processing network, please refer to [link / reference]. Figure 2 The implementation method shown in step S22 will not be described again here.
[0148] like Figure 12 As shown, for each training sample, the feature alignment network determines the offset feature value of each pixel of the initial image feature corresponding to the training sample relative to the pixel at the same position in the corresponding standard image feature. Based on the initial image feature corresponding to the training sample and the offset feature value corresponding to each pixel of the initial image feature, the first fused image feature is determined. For details on implementing the feature processing of the aforementioned flexible alignment module, please refer to [link to relevant documentation]. Figure 2 The implementation method shown in step S23 will not be described again here.
[0149] like Figure 12As shown, for each training sample, the type prediction network predicts the object type of the sample object in the training sample based on the first fused image features corresponding to that training sample. See [link to documentation] for details. Figure 2 The implementation method shown in step S24 will not be described again here.
[0150] Step S113: Determine the actual object type of the sample objects in each training sample group, determine the total training loss based on the actual object type of the sample objects in each training sample group and the predicted object type, iteratively train the initial model based on the training sample set until the total training loss meets the training termination condition, and stop training, and determine the model at the time of training termination as the defect detection model.
[0151] In some feasible implementations, during the process of determining each set of training samples, if both images in the training sample are standard images of a standard object, then the actual object type to which the sample object in the training sample belongs is determined to be the second type; otherwise, the actual object type to which the sample object in the training sample belongs is determined to be the first type.
[0152] Furthermore, during the training of the initial model, the total training loss can be determined based on the actual object type and the predicted object type of the sample objects in each training sample set. Based on this, training can be stopped when the total training loss meets the training termination condition, and the model at the point of training stoppage can be determined as the final defect detection model.
[0153] The training termination condition mentioned above can be either the total training loss reaching convergence or the total training loss being less than the loss threshold. The specific condition can be determined based on the actual application scenario requirements and is not restricted here.
[0154] The total training loss mentioned above can be determined by cross-entropy loss or other loss functions, and there are no restrictions on this.
[0155] In this embodiment, the training method and defect detection method of the aforementioned defect detection model can be implemented based on machine learning technology in artificial intelligence. Artificial intelligence is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. Artificial intelligence studies the design principles and implementation methods of various intelligent machines, enabling them to have perception, reasoning, and decision-making functions. Machine learning specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills, reorganize existing knowledge structures to continuously improve their performance, and thus realize the training process of the aforementioned Hidden Markov Model.
[0156] In this application embodiment, the feature processing process can be implemented based on cloud computing technology in the field of cloud technology to improve computing efficiency.
[0157] Cloud technology refers to a hosting technology that unifies hardware, software, and network resources within a wide area network (WAN) or local area network (LAN) to achieve data computation, storage, processing, and sharing. Cloud computing is a computing model, a product of the convergence of traditional computer and network technologies such as grid computing, distributed computing, parallel computing, utility computing, network storage technologies, virtualization, and load balancing. Cloud computing distributes computing tasks across a resource pool composed of numerous computers, enabling various application systems to obtain computing power, storage space, and information services as needed. The network providing these resources is called the "cloud." Resources in the "cloud" are infinitely scalable, readily available, on-demand, expandable, and pay-as-you-go.
[0158] See Figure 13 , Figure 13 This is a schematic diagram of the defect detection device provided in an embodiment of this application. The defect detection device provided in an embodiment of this application includes:
[0159] Image determination module 131 is used to determine an initial image of the object to be detected and a standard image of a standard object, wherein the initial image and the standard image are obtained based on the same shooting conditions;
[0160] The feature extraction module 132 is used to determine the initial image features of the initial image and the standard image features of the standard image.
[0161] Feature processing module 133 is used to determine the offset feature value of each pixel of the initial image feature relative to the pixel at the same position in the standard image feature, and to determine the first fused image feature based on the initial image feature and the offset feature value corresponding to each pixel of the initial image feature.
[0162] The type determination module 134 is used to determine the object type of the object to be detected based on the first fused image features. The object type includes a first type and a second type. When the object to be detected belongs to the first type, the object to be detected has defects compared to the standard object. When the object to be detected belongs to the first type, the object to be detected does not have defects compared to the standard object.
[0163] In some feasible implementations, the feature extraction module 132 described above is used for:
[0164] The initial image features of the above initial image are determined based on the first feature extraction network;
[0165] The standard image features of the above standard images are determined based on the second feature extraction network;
[0166] The first feature extraction network and the second feature extraction network have the same network structure. The network weights of the first feature extraction network are determined based on the network weights of the second feature extraction network and preset weight components.
[0167] In some feasible implementations, the feature processing module 133 described above is used for:
[0168] Based on the above initial image features and the above standard image features, determine the offset position of each pixel in the above initial image features relative to the pixel at the same position in the above standard image features;
[0169] Determine the target feature value corresponding to the offset position of each pixel of the initial image feature in the standard image feature, and determine the target feature value corresponding to each pixel of the initial image feature as the offset feature value relative to the pixel at the same position in the standard image feature.
[0170] In some feasible implementations, the feature processing module 133 described above is used for:
[0171] The second fused image features are determined based on the aforementioned initial image features and the aforementioned standard image features;
[0172] Based on the second fused image features, the offset information of each pixel of the initial image features is determined relative to the pixel at the same position in the standard image features. The offset information of each pixel of the initial image features includes the horizontal offset, the horizontal offset direction, the vertical offset, and the vertical offset direction.
[0173] For each pixel of the initial image features described above, the offset position of the pixel relative to the pixel at the same position in the standard image is determined based on the offset information of that pixel.
[0174] In some feasible implementations, the feature processing module 133 described above is used for:
[0175] Based on the aforementioned second fused image features, a first offset feature and a second offset feature are determined. Each pixel of the first offset feature is used to characterize the lateral offset reference value of the pixel at the same position in the aforementioned initial image features, and each pixel of the second offset feature is used to characterize the vertical offset reference value of the pixel at the same position in the aforementioned initial image features.
[0176] For each pixel of the initial image feature, the lateral offset and lateral offset direction of the pixel relative to the pixel at the same position in the standard image feature are determined based on the lateral offset reference value of the pixel. The longitudinal offset and longitudinal offset direction of the pixel relative to the pixel at the same position in the standard image feature are determined based on the longitudinal offset reference value of the pixel.
[0177] In some feasible implementations, for each pixel of the initial image features described above, the feature processing module 133 is configured to:
[0178] If the lateral offset reference value of the pixel is less than the first threshold, the lateral offset direction of the pixel is determined to be to the left. If the lateral offset reference value of the pixel is greater than the first threshold, the lateral offset direction of the pixel is determined to be to the right. The lateral offset amount of the pixel is determined based on the lateral offset reference value of the pixel, the first threshold, and the lateral offset coefficient.
[0179] The aforementioned feature processing module 133 is used for:
[0180] If the vertical offset reference value of the pixel is less than the second threshold, the vertical offset direction of the pixel is determined to be upward. If the vertical offset reference value of the pixel is greater than the second threshold, the vertical offset direction of the pixel is determined to be downward. The vertical offset amount of the pixel is determined based on the vertical offset reference value, the second threshold, and the vertical offset coefficient.
[0181] In some feasible implementations, the feature processing module 133 described above is used for:
[0182] The feature value of each pixel in the initial image features and the corresponding offset feature value are fused to obtain the third fused image feature;
[0183] The first fused image features are determined based on the third fused image features and the initial image features.
[0184] In some feasible implementations, the type determination module 134 described above is further configured to:
[0185] In response to the fact that the object to be detected belongs to the first type, the target defects in the object to be detected are determined based on the initial image of the object to be detected, and the target defects are classified.
[0186] In some feasible implementations, the type determination module 134 described above is further configured to:
[0187] Determine the heatmap of the initial image of the object to be detected, and perform binarization processing on the heatmap to obtain the target grayscale image;
[0188] Each connected component in the target grayscale image is determined, and the target defect in the target object is determined from the image region corresponding to each connected component in the initial image of the target object.
[0189] In some feasible implementations, determining the initial image features and the standard image features, determining the offset feature value of each pixel of the initial image features, determining the first fused image features, and determining the object type of the object to be detected are based on a defect detection model. This defect detection model is trained using a training module, which is used to:
[0190] The training sample set is determined, which includes multiple sets of training samples. Each set of training samples includes an initial image of a sample object and a standard image of the corresponding standard object.
[0191] Each set of the above training samples is input into the initial model to predict the object type of the sample objects in each set of the above training samples based on the initial model.
[0192] Determine the actual object type of the sample objects in each group of the above training samples. Based on the actual object type of the sample objects in each group of the above training samples and the predicted object type, determine the total training loss. Iterate the training of the initial model based on the training sample set until the total training loss meets the training termination condition, and stop training. The model at the time of training termination is determined as the above defect detection model.
[0193] In specific implementation, the above-mentioned device can perform the above-described actions through its built-in functional modules. Figure 2 and / or Figure 11 The implementation methods provided for each step are detailed in the above-mentioned implementation methods, and will not be repeated here.
[0194] See Figure 14 , Figure 14 This is a schematic diagram of the structure of the electronic device provided in an embodiment of this application. For example... Figure 14 As shown, the electronic device 1400 in this embodiment may include: a processor 1401, a network interface 1404, and a memory 1405. Furthermore, the electronic device 1400 may also include: an object interface 1403, and at least one communication bus 1402. The communication bus 1402 is used to implement communication between these components. The object interface 1403 may include a display screen and a keyboard; optionally, the object interface 1403 may also include a standard wired interface or a wireless interface. The network interface 1404 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1405 may be a high-speed RAM or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory 1405 may also be at least one storage device located remotely from the aforementioned processor 1401. Figure 14 As shown, the memory 1405, which is a computer-readable storage medium, may include an operating system, a network communication module, an object interface module, and a device control application.
[0195] exist Figure 14 In the illustrated electronic device 1400, the network interface 1404 provides network communication functionality; the object interface 1403 is primarily used to provide an input interface for objects; and the processor 1401 can be used to call the device control application stored in the memory 1405 to achieve:
[0196] An initial image of the object to be detected and a standard image of a standard object are determined, wherein the initial image and the standard image are obtained based on the same shooting conditions;
[0197] Determine the initial image features of the initial image and the standard image features of the standard image.
[0198] Determine the offset feature value of each pixel of the initial image feature relative to the pixel at the same position in the standard image feature, and determine the first fused image feature based on the initial image feature and the offset feature value corresponding to each pixel of the initial image feature;
[0199] Based on the first fused image features, the object type of the object to be detected is determined. The object type includes a first type and a second type. When the object to be detected belongs to the first type, the object to be detected has defects compared to the standard object. When the object to be detected belongs to the first type, the object to be detected does not have defects compared to the standard object.
[0200] In some feasible implementations, the processor 1401 described above is used for:
[0201] The initial image features of the above initial image are determined based on the first feature extraction network;
[0202] The standard image features of the above standard images are determined based on the second feature extraction network;
[0203] The first feature extraction network and the second feature extraction network have the same network structure. The network weights of the first feature extraction network are determined based on the network weights of the second feature extraction network and preset weight components.
[0204] In some feasible implementations, the processor 1401 described above is used for:
[0205] Based on the above initial image features and the above standard image features, determine the offset position of each pixel in the above initial image features relative to the pixel at the same position in the above standard image features;
[0206] Determine the target feature value corresponding to the offset position of each pixel of the initial image feature in the standard image feature, and determine the target feature value corresponding to each pixel of the initial image feature as the offset feature value relative to the pixel at the same position in the standard image feature.
[0207] In some feasible implementations, the processor 1401 described above is used for:
[0208] The second fused image features are determined based on the aforementioned initial image features and the aforementioned standard image features;
[0209] Based on the second fused image features, the offset information of each pixel of the initial image features is determined relative to the pixel at the same position in the standard image features. The offset information of each pixel of the initial image features includes the horizontal offset, the horizontal offset direction, the vertical offset, and the vertical offset direction.
[0210] For each pixel of the initial image features described above, the offset position of the pixel relative to the pixel at the same position in the standard image is determined based on the offset information of that pixel.
[0211] In some feasible implementations, the processor 1401 described above is used for:
[0212] Based on the aforementioned second fused image features, a first offset feature and a second offset feature are determined. Each pixel of the first offset feature is used to characterize the lateral offset reference value of the pixel at the same position in the aforementioned initial image features, and each pixel of the second offset feature is used to characterize the vertical offset reference value of the pixel at the same position in the aforementioned initial image features.
[0213] For each pixel of the initial image feature, the lateral offset and lateral offset direction of the pixel relative to the pixel at the same position in the standard image feature are determined based on the lateral offset reference value of the pixel. The longitudinal offset and longitudinal offset direction of the pixel relative to the pixel at the same position in the standard image feature are determined based on the longitudinal offset reference value of the pixel.
[0214] In some feasible implementations, for each pixel of the initial image features described above, the processor 1401 is used to:
[0215] If the lateral offset reference value of the pixel is less than the first threshold, the lateral offset direction of the pixel is determined to be to the left. If the lateral offset reference value of the pixel is greater than the first threshold, the lateral offset direction of the pixel is determined to be to the right. The lateral offset amount of the pixel is determined based on the lateral offset reference value of the pixel, the first threshold, and the lateral offset coefficient.
[0216] The determination of the vertical offset amount and direction of the pixel relative to the pixel at the same position in the standard image features, based on the vertical offset reference value of the pixel, includes:
[0217] If the vertical offset reference value of the pixel is less than the second threshold, the vertical offset direction of the pixel is determined to be upward. If the vertical offset reference value of the pixel is greater than the second threshold, the vertical offset direction of the pixel is determined to be downward. The vertical offset amount of the pixel is determined based on the vertical offset reference value, the second threshold, and the vertical offset coefficient.
[0218] In some feasible implementations, the processor 1401 described above is used for:
[0219] The feature value of each pixel in the initial image features and the corresponding offset feature value are fused to obtain the third fused image feature;
[0220] The first fused image features are determined based on the third fused image features and the initial image features.
[0221] In some feasible implementations, the processor 1401 is further configured to:
[0222] In response to the fact that the object to be detected belongs to the first type, the target defects in the object to be detected are determined based on the initial image of the object to be detected, and the target defects are classified.
[0223] In some feasible implementations, the processor 1401 described above is used for:
[0224] Determine the heatmap of the initial image of the object to be detected, and perform binarization processing on the heatmap to obtain the target grayscale image;
[0225] Each connected component in the target grayscale image is determined, and the target defect in the target object is determined from the image region corresponding to each connected component in the initial image of the target object.
[0226] In some feasible implementations, determining the initial image features and the standard image features, determining the offset feature value of each pixel of the initial image features, determining the first fused image features, and determining the object type of the object to be detected are based on a defect detection model, which is trained by the processor 1401 in the following manner:
[0227] The training sample set is determined, which includes multiple sets of training samples. Each set of training samples includes an initial image of a sample object and a standard image of the corresponding standard object.
[0228] Each set of the above training samples is input into the initial model to predict the object type of the sample objects in each set of the above training samples based on the initial model.
[0229] Determine the actual object type of the sample objects in each group of the above training samples. Based on the actual object type of the sample objects in each group of the above training samples and the predicted object type, determine the total training loss. Iterate the training of the initial model based on the training sample set until the total training loss meets the training termination condition, and stop training. The model at the time of training termination is determined as the above defect detection model.
[0230] It should be understood that in some feasible implementations, the processor 1401 described above may be a central processing unit (CPU), which may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor. The memory may include read-only memory and random access memory, and provides instructions and data to the processor. A portion of the memory may also include non-volatile random access memory. For example, the memory may also store device type information.
[0231] In specific implementation, the aforementioned electronic device 1400 can perform the above-described actions through its built-in functional modules. Figure 2 and / or Figure 11 The implementation methods provided for each step are detailed in the above-mentioned implementation methods, and will not be repeated here.
[0232] This application also provides a computer-readable storage medium storing a computer program that is executed by a processor to implement... Figure 2 and / or Figure 11 The methods provided in each step are detailed in the implementation methods provided in the above steps, and will not be repeated here.
[0233] The aforementioned computer-readable storage medium can be an internal storage unit of the apparatus or electronic device provided in any of the foregoing embodiments, such as a hard disk or memory of the electronic device. The computer-readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the electronic device. The aforementioned computer-readable storage medium can also include magnetic disks, optical disks, read-only memory (ROM), or random access memory (RAM), etc. Furthermore, the computer-readable storage medium can include both internal storage units and external storage devices of the electronic device. The computer-readable storage medium is used to store the computer program and other programs and data required by the electronic device. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.
[0234] This application provides a computer program product, which includes a computer program that is executed by a processor. Figure 2 and / or Figure 11 The methods provided for each step in the process.
[0235] The terms "first," "second," etc., used in the claims, description, and drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or electronic device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or electronic devices. References to "embodiment" herein mean that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The presentation of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments. The term "and / or" as used in this application's description and appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0236] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Those skilled in the art can implement the described functions using different methods for each specific application, but such implementations should not be considered beyond the scope of this application.
[0237] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application shall still fall within the scope of this application.
Claims
1. A defect detection method, characterized in that, The method includes: An initial image of the object to be detected and a standard image of a standard object are determined, wherein the initial image and the standard image are obtained based on the same shooting conditions; Determine the initial image features of the initial image and the standard image features of the standard image; The offset feature value of each pixel of the initial image feature is determined relative to the pixel at the same position in the standard image feature. Based on the initial image feature and the offset feature value corresponding to each pixel of the initial image feature, the initial image feature is aligned with the standard image feature to determine a first fused image feature. The first fused image feature is used to characterize the feature information corresponding to the initial image feature after eliminating the deformation feature differences with the standard image feature. Based on the first fused image features, the object type to which the object to be detected belongs is determined. The object type includes a first type and a second type. When the object to be detected belongs to the first type, the object to be detected has defects compared to the standard object. When the object to be detected belongs to the first type, the object to be detected does not have defects compared to the standard object.
2. The method according to claim 1, characterized in that, The process of determining the initial image features of the initial image and the standard image features of the standard image includes: The initial image features of the initial image are determined based on the first feature extraction network; The standard image features of the standard image are determined based on the second feature extraction network; The first feature extraction network and the second feature extraction network have the same network structure. The network weights of the first feature extraction network are determined based on the network weights of the second feature extraction network and preset weight components.
3. The method according to claim 1, characterized in that, The step of determining the offset feature value of each pixel in the initial image feature relative to the pixel at the same position in the standard image feature includes: Based on the initial image features and the standard image features, determine the offset position of each pixel in the initial image features relative to the pixel at the same position in the standard image features; The target feature value corresponding to the offset position of each pixel of the initial image feature in the standard image feature is determined, and the target feature value corresponding to each pixel of the initial image feature is determined as the offset feature value relative to the pixel at the same position in the standard image feature.
4. The method according to claim 3, characterized in that, The step of determining the offset position of each pixel in the initial image features relative to the pixel at the same position in the standard image features, based on the initial image features and the standard image features, includes: The second fused image features are determined based on the initial image features and the standard image features; Based on the second fused image features, the offset information of each pixel of the initial image feature relative to the pixel at the same position in the standard image feature is determined. The offset information of each pixel of the initial image feature includes the horizontal offset, the horizontal offset direction, the vertical offset, and the vertical offset direction. For each pixel of the initial image feature, the offset position of the pixel relative to the pixel at the same position in the standard image is determined based on the offset information of the pixel.
5. The method according to claim 4, characterized in that, The step of determining the offset information of each pixel in the initial image feature relative to the pixel at the same position in the standard image feature based on the second fused image feature includes: Based on the second fused image features, a first offset feature and a second offset feature are determined. Each pixel of the first offset feature is used to characterize the horizontal offset reference value of the pixel at the same position in the initial image features, and each pixel of the second offset feature is used to characterize the vertical offset reference value of the pixel at the same position in the initial image features. For each pixel of the initial image feature, the lateral offset and lateral offset direction of the pixel relative to the pixel at the same position in the standard image feature are determined based on the lateral offset reference value of the pixel. The longitudinal offset and longitudinal offset direction of the pixel relative to the pixel at the same position in the standard image feature are determined based on the longitudinal offset reference value of the pixel.
6. The method according to claim 5, characterized in that, For each pixel of the initial image feature, determining the lateral offset and direction of that pixel relative to a pixel at the same position in the standard image feature based on the lateral offset reference value of that pixel includes: In response to the lateral offset reference value of the pixel being less than the first threshold, the lateral offset direction of the pixel is determined to be to the left; in response to the lateral offset reference value of the pixel being greater than the first threshold, the lateral offset direction of the pixel is determined to be to the right; and the lateral offset amount of the pixel is determined based on the lateral offset reference value of the pixel, the first threshold, and the lateral offset coefficient. The determination of the vertical offset amount and direction of the pixel relative to the pixel at the same position in the standard image feature based on the vertical offset reference value of the pixel includes: If the vertical offset reference value of the pixel is less than the second threshold, the vertical offset direction of the pixel is determined to be upward. If the vertical offset reference value of the pixel is greater than the second threshold, the vertical offset direction of the pixel is determined to be downward. The vertical offset amount of the pixel is determined based on the vertical offset reference value, the second threshold, and the vertical offset coefficient.
7. The method according to claim 1, characterized in that, The step of aligning the initial image features with the standard image features based on the initial image features and the offset feature values corresponding to each pixel of the initial image features to determine the first fused image features includes: The feature value of each pixel of the initial image feature and the corresponding offset feature value are fused to align the initial image feature with the standard image feature, thereby obtaining the third fused image feature; The first fused image feature is determined based on the third fused image feature and the initial image feature.
8. The method according to claim 1, characterized in that, The method further includes: In response to the object to be detected belonging to the first type, target defects in the object to be detected are determined based on the initial image of the object to be detected, and the target defects are classified.
9. The method according to claim 8, characterized in that, The step of determining the target defect in the object to be detected based on the initial image of the object to be detected includes: Determine the heatmap of the initial image of the object to be detected, and perform binarization processing on the heatmap to obtain the target grayscale image; Each connected component in the target grayscale image is determined, and the target defect in the object to be detected is determined from the image region corresponding to each connected component in the initial image of the object to be detected.
10. The method according to claim 1, characterized in that, The determination of the initial image features and the standard image features, the determination of the offset feature value of each pixel of the initial image features, the determination of the first fused image features, and the determination of the object type to which the object to be detected belongs are implemented based on a defect detection model, which is trained in the following manner: A training sample set is determined, comprising multiple sets of training samples. Each set of training samples includes an initial image of a sample object and a standard image of the corresponding standard object. Each set of training samples is input into the initial model to predict the object type of the sample objects in each set of training samples based on the initial model. The actual object type of the sample objects in each group of training samples is determined. The total training loss is determined based on the actual object type of the sample objects in each group of training samples and the predicted object type. The initial model is iteratively trained based on the training sample set until the total training loss meets the training termination condition. The model at the time of training termination is then determined as the defect detection model.
11. A defect detection device, characterized in that, The device includes: An image determination module is used to determine an initial image of the object to be detected and a standard image of a standard object, wherein the initial image and the standard image are obtained based on the same shooting conditions; The feature extraction module is used to determine the initial image features of the initial image and the standard image features of the standard image; The feature processing module is used to determine the offset feature value of each pixel of the initial image feature relative to the pixel at the same position in the standard image feature, and based on the initial image feature and the offset feature value corresponding to each pixel of the initial image feature, to perform feature alignment of the initial image feature to the standard image feature, and to determine a first fused image feature. The first fused image feature is used to characterize the feature information corresponding to the initial image feature after eliminating the deformation feature differences with the standard image feature. The type determination module is used to determine the object type to which the object to be detected belongs based on the first fused image features. The object type includes a first type and a second type. When the object to be detected belongs to the first type, the object to be detected has defects compared to the standard object. When the object to be detected belongs to the first type, the object to be detected does not have defects compared to the standard object.
12. The apparatus according to claim 11, characterized in that, The feature extraction module is used for: The initial image features of the initial image are determined based on the first feature extraction network; The standard image features of the standard image are determined based on the second feature extraction network; The first feature extraction network and the second feature extraction network have the same network structure. The network weights of the first feature extraction network are determined based on the network weights of the second feature extraction network and preset weight components.
13. The apparatus according to claim 11, characterized in that, The feature processing module is used for: Based on the initial image features and the standard image features, determine the offset position of each pixel in the initial image features relative to the pixel at the same position in the standard image features; The target feature value corresponding to the offset position of each pixel of the initial image feature in the standard image feature is determined, and the target feature value corresponding to each pixel of the initial image feature is determined as the offset feature value relative to the pixel at the same position in the standard image feature.
14. The apparatus according to claim 13, characterized in that, The feature processing module is used for: The second fused image features are determined based on the initial image features and the standard image features; Based on the second fused image features, the offset information of each pixel of the initial image feature relative to the pixel at the same position in the standard image feature is determined. The offset information of each pixel of the initial image feature includes the horizontal offset, the horizontal offset direction, the vertical offset, and the vertical offset direction. For each pixel of the initial image feature, the offset position of the pixel relative to the pixel at the same position in the standard image is determined based on the offset information of the pixel.
15. The apparatus according to claim 14, characterized in that, The feature processing module is used for: Based on the second fused image features, a first offset feature and a second offset feature are determined. Each pixel of the first offset feature is used to characterize the horizontal offset reference value of the pixel at the same position in the initial image features, and each pixel of the second offset feature is used to characterize the vertical offset reference value of the pixel at the same position in the initial image features. For each pixel of the initial image feature, the lateral offset and lateral offset direction of the pixel relative to the pixel at the same position in the standard image feature are determined based on the lateral offset reference value of the pixel. The longitudinal offset and longitudinal offset direction of the pixel relative to the pixel at the same position in the standard image feature are determined based on the longitudinal offset reference value of the pixel.
16. The apparatus according to claim 15, characterized in that, For each pixel of the initial image features, the feature processing module is configured to: In response to the lateral offset reference value of the pixel being less than the first threshold, the lateral offset direction of the pixel is determined to be to the left; in response to the lateral offset reference value of the pixel being greater than the first threshold, the lateral offset direction of the pixel is determined to be to the right; and the lateral offset amount of the pixel is determined based on the lateral offset reference value of the pixel, the first threshold, and the lateral offset coefficient. The determination of the vertical offset amount and direction of the pixel relative to the pixel at the same position in the standard image feature based on the vertical offset reference value of the pixel includes: If the vertical offset reference value of the pixel is less than the second threshold, the vertical offset direction of the pixel is determined to be upward. If the vertical offset reference value of the pixel is greater than the second threshold, the vertical offset direction of the pixel is determined to be downward. The vertical offset amount of the pixel is determined based on the vertical offset reference value, the second threshold, and the vertical offset coefficient.
17. The apparatus according to claim 11, characterized in that, The feature processing module is configured to include: The feature value of each pixel of the initial image feature and the corresponding offset feature value are fused to align the initial image feature with the standard image feature, thereby obtaining the third fused image feature; The first fused image feature is determined based on the third fused image feature and the initial image feature.
18. The apparatus according to claim 11, characterized in that, The type determination module is also used for: In response to the object to be detected belonging to the first type, target defects in the object to be detected are determined based on the initial image of the object to be detected, and the target defects are classified.
19. The apparatus according to claim 18, characterized in that, The type determination module is used for: Determine the heatmap of the initial image of the object to be detected, and perform binarization processing on the heatmap to obtain the target grayscale image; Each connected component in the target grayscale image is determined, and the target defect in the object to be detected is determined from the image region corresponding to each connected component in the initial image of the object to be detected.
20. The apparatus according to claim 11, characterized in that, Determining the initial image features and the standard image features, determining the offset feature value of each pixel of the initial image features, determining the first fused image features, and determining the object type to which the object to be detected belongs are implemented based on a defect detection model. This defect detection model is trained using a training module, which is used for: A training sample set is determined, comprising multiple sets of training samples. Each set of training samples includes an initial image of a sample object and a standard image of the corresponding standard object. Each set of training samples is input into the initial model to predict the object type of the sample objects in each set of training samples based on the initial model. The actual object type of the sample objects in each group of training samples is determined. The total training loss is determined based on the actual object type of the sample objects in each group of training samples and the predicted object type. The initial model is iteratively trained based on the training sample set until the total training loss meets the training termination condition. The model at the time of training termination is then determined as the defect detection model.
21. An electronic device, characterized in that, It includes a processor and a memory, which are interconnected; The memory is used to store computer programs; The processor is configured to perform the method as described in any one of claims 1 to 10 when the computer program is invoked.
22. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that is executed by a processor to implement the method of any one of claims 1 to 10.
23. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Grad-CAM algorithm-based high-speed rail contact network insulator damage accurate positioning method
CN110176001A
Multi-scene power transformation equipment defect change detection method based on optical flow feature fusion
CN114066864A
Detection model training method and device, defect detection method and device and storage medium
CN115311273A