Alloy surface machining defect detection method, device, equipment, medium and product
By locally segmenting and identifying defects in alloy surface images based on a semantic segmentation network model that segments everything, the problems of low efficiency and poor accuracy in detecting defects in alloy surface processing are solved, and real-time monitoring and efficient quality control are achieved.
Patent Information
- Application Number
- CN202510098568.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-10-17
AI Technical Summary
The existing technology has low efficiency and poor accuracy in detecting defects on alloy surface processing, cannot monitor in real time, has high labor costs, and has poor reliability in alloy quality control.
A semantic segmentation network model based on segmentation of everything is adopted to collect alloy surface images and divide them into local blocks, identify and mark defect areas, and determine the degree of defects based on the proportion of defect areas.
It improves the efficiency and accuracy of alloy surface processing defect detection, realizes real-time monitoring of the processing process, reduces labor costs, and improves the reliability of quality control.
Smart Images

Figure CN120807383A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer image segmentation, and in particular to a method and device for detecting defects on an alloy surface, an apparatus, a medium and a product. BACKGROUND
[0002] Alloys are widely used in industries, construction and aerospace, etc. In order to improve the surface hardness and corrosion resistance of alloys, surface processing is required. However, if the process parameters are not properly controlled during surface processing, defects may occur on the alloy surface, reducing the protective performance and service life of the alloy. Therefore, timely detection and evaluation of alloy surface processing defects are crucial to ensuring product quality.
[0003] In the prior art, alloy surface processing defect detection requires professional personnel to measure on site. Due to the limited observation ability of the human eye, defects that are extremely small, hidden in location or located deep inside the alloy are difficult to detect, resulting in insufficient detection accuracy. These defects, although small, can be a key factor affecting the performance and safety of the alloy. For example, if there are small cracks or scratches on the alloy surface, they may cause component fatigue fracture in extreme environments, leading to serious safety accidents. In addition, manual detection is usually performed in stages and cannot achieve continuous and real-time monitoring of the entire production process, which may result in missed detection of defects occurring during production. As production pace accelerates and production scale expands, the efficiency and coverage of manual detection are difficult to meet the needs of real-time monitoring, thereby increasing the risk of product quality and reducing the reliability of alloy quality control. SUMMARY
[0004] The present application provides a method and device for detecting defects on an alloy surface, an apparatus, a medium and a product to solve the problems of low efficiency, poor accuracy, inability to monitor the alloy processing process in real time, high labor cost and poor reliability of alloy quality control when detecting defects on the alloy surface.
[0005] According to an aspect of an embodiment of the present application, a method for detecting defects on an alloy surface is provided, comprising:
[0006] Collecting a target alloy image corresponding to a target alloy after processing, and dividing the target alloy image into at least one target local patch according to the input data requirements of a pre-trained network model;
[0007] Inputting each target local patch into the pre-trained network model to obtain an abnormal defect feature map corresponding to each target local patch; wherein the pre-trained network model uses a semantic segmentation network based on segmentation of everything;
[0008] According to the position coordinates of the abnormal points in each abnormal defect feature map and the segmentation position of each target local block belonging to the abnormal defect feature map in the target alloy image, a defect region is marked in the target alloy image.
[0009] According to the proportion of the defect region in the target alloy image, the surface processing defect degree of the target alloy after processing is determined.
[0010] According to another aspect of the embodiment of the present application, a detection device for alloy surface processing defects is also provided, comprising:
[0011] An image segmentation module is configured to collect a target alloy image corresponding to a target alloy after processing, and segment the target alloy image into at least one target local block according to the input data requirement of a pre-trained network model.
[0012] A feature map acquisition module is configured to input each target local block into the pre-trained network model to obtain an abnormal defect feature map corresponding to each target local block, wherein the pre-trained network model adopts a semantic segmentation network based on segmentation of all.
[0013] A defect marking module is configured to mark a defect region in the target alloy image according to the position coordinates of the abnormal points in each abnormal defect feature map and the segmentation position of each target local block belonging to the abnormal defect feature map in the target alloy image.
[0014] A degree determination module is configured to determine the surface processing defect degree of the target alloy after processing according to the proportion of the defect region in the target alloy image.
[0015] According to another aspect of the embodiment of the present application, an electronic device is also provided, comprising:
[0016] At least one processor; and
[0017] A memory in communication connection with the at least one processor; wherein,
[0018] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the detection method for alloy surface processing defects according to any one of the embodiments of the present application.
[0019] According to another aspect of the embodiment of the present application, a computer readable storage medium is also provided, which stores computer instructions for enabling a processor to execute the detection method for alloy surface processing defects according to any one of the embodiments of the present application.
[0020] According to another aspect of the embodiments of the present application, there is also provided a computer program product comprising a computer program which, when executed by a processor, implements the steps of the method according to any of the embodiments of the present application.
[0021] The technical solution of the embodiments of the present application acquires the target alloy image after processing, and according to the input data requirements of the pre-trained network model, the image is cut into target local blocks and input into the model to obtain the feature maps of each abnormal defect; according to the position coordinates of the abnormal points in each image and the cutting position of each target local block in the target alloy image, the defect area is marked; according to the proportion of the defect area in the target alloy image, the surface processing defect degree of the target alloy after processing is determined. The pre-trained network model is used to obtain the feature maps of the abnormal defects of the target alloy image, and the defect degree is quantified according to the proportion of the defect area in the target alloy image, which improves the detection efficiency and accuracy of the alloy surface processing defects, and can monitor the alloy processing process in real time, reduces the labor cost, and improves the reliability of the alloy quality control.
[0022] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0024] Figure 1 is a flowchart of a method for detecting alloy surface processing defects according to an embodiment of the present application;
[0025] Figure 2 is a flowchart of another method for detecting alloy surface processing defects according to an embodiment of the present application;
[0026] Figure 3 is a flowchart of a method for detecting and evaluating alloy surface processing defects according to an embodiment of the present application;
[0027] Figure 4 is a network structure diagram of a large-scale pre-training adaptive network according to an embodiment of the present application;
[0028] Figure 5 is a structure diagram of a device for detecting alloy surface processing defects according to an embodiment of the present application;
[0029] Figure 6 It is a structural schematic diagram of an electronic device for implementing the method for detecting alloy surface processing defects according to an embodiment of the present invention. DETAILED DESCRIPTION
[0030] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0031] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0032] Example 1
[0033] Figure 1 This is a flow chart of a method for detecting alloy surface processing defects provided in the first embodiment of the present invention. This embodiment is applicable to detecting alloy surface processing defects by collecting images of the processed alloy surface. This method can be performed by a device for detecting alloy surface processing defects. The device can be implemented in the form of hardware and / or software and can generally be configured in an electronic device. Figure 1 As shown, the method includes:
[0034] S110 , collecting a target alloy image corresponding to the processed target alloy, and dividing the target alloy image into at least one target local image block according to the input data requirements of the pre-trained network model.
[0035] In the embodiments of the present application, the target alloy can be specifically understood as an alloy material that needs to be detected and analyzed for surface processing defects in the application scenarios of alloy surface processing such as anodic oxidation and plating. The target alloy image can be specifically understood as an image reflecting the characteristics of the target alloy surface and defects, including alloy surface texture, color, shape, and defects (such as scratches, pits, and cracks, etc.), obtained by image acquisition devices (such as industrial cameras and cameras, etc.). For example, in the target alloy image after anodic oxidation treatment on the surface of an aluminum alloy, the thickness, uniformity, and pore characteristics of the oxide layer can be obtained to judge the effect and quality of anodic oxidation. Typically, the target alloy can be described by a rectangular image uniquely determined by four vertices. The input data requirement can be specifically understood as the requirement of the pre-trained network model for the form of input data when performing image processing and analysis. For example, the model can need to input images with specific size, number of color channels, and data type, etc.
[0036] Since the size of the target alloy image exceeds the input data requirement of the pre-trained network model, or the information contained in the image is relatively large, directly inputting the entire image can cause excessive computational burden of the model or poor segmentation effect, etc., and the target alloy image needs to be cut into at least one target local patch. The target local patch can be specifically understood as an image block containing local area information cut from the target alloy image. Each target local patch corresponds to a specific area in the target alloy image and can be independently input into the pre-trained network model for processing and analysis.
[0037] When the defects in the target alloy image are relatively evenly distributed and the defect size is relatively small, the target alloy image can be uniformly cut according to the fixed size of the grid according to the input data requirement of the pre-trained network model, and each grid corresponds to a target local patch. For example, the image can be cut into multiple 64x64 pixel local patches; when the defects in the target alloy image are unevenly distributed or the defect size difference is large, a dynamic size sliding window can be used to slide on the target alloy image according to the input data requirement of the pre-trained network model, and each time the sliding window size step, the area covered by the window is taken as a target local patch. The size and step of the sliding window are adjusted to adapt to different defect characteristics, realizing continuous cutting of the target alloy image. For example, in the area where the defects are relatively dense, a smaller sliding window can be used to improve the detection accuracy of small size defects and avoid merging multiple defects together due to the window being too large. In the area where the defects are relatively sparse, the size of the sliding window can be increased to reduce the amount of calculation and processing time. At the same time, a larger window can capture more background information in a larger range, providing more abundant context for defect recognition.
[0038] S120, input each target local patch into the pre-trained network model to obtain an abnormal defect feature map corresponding to each target local patch.
[0039] The pre-trained network model adopts a semantic segmentation network based on segment anything.
[0040] In the embodiment of the application, the pre-trained network model can be specifically understood as follows: by being trained on a large number of alloy surface processing defect image data, the pre-trained network model has certain feature extraction and image processing capabilities, can recognize and understand various features and patterns in the image, and can be used as a basic model for extracting feature information in the target local patch. By inputting the target local patch into the pre-trained network model, the model can identify the abnormal defect features in the patch and generate a corresponding abnormal defect feature map. The abnormal defect feature map can be specifically understood as a feature representation map output by the network model, which contains feature information of the abnormal defect in the target local patch, such as the shape, size, position and texture of the defect. The image is usually presented in the form of a binary image, in which the pixel value of the defect area is different from that of the background area, so as to distinguish the abnormal defect area in the local patch. The semantic segmentation network based on segment anything (Segment Anything Model, SAM) is an image segmentation model used to identify and segment the abnormal defect area in the target local patch and generate a high-quality abnormal defect feature map.
[0041] S130, according to the position coordinates of the abnormal points in each abnormal defect feature map and the segmentation position of each target local patch in the target alloy image, marking the defect area in the target alloy image.
[0042] Specifically, after obtaining each abnormal defect feature map, the position coordinates of the abnormal points are extracted from the abnormal defect feature map by image processing techniques (such as binarization or edge detection, etc.). For example, assuming that the pixel value of the defect area in the abnormal defect feature map is 255 and the pixel value of the background area is 0, the coordinates of the pixel value 255 can be recorded by traversing each pixel of the feature map, so as to obtain the relative position coordinates of the abnormal points in the target local patch.
[0043] When the target alloy image is segmented into target local patches, the position information of each local patch in the target alloy image can be recorded by calculating the starting coordinate point (such as the top-left corner coordinate) of the local patch. For example, if the target alloy image is uniformly segmented into multiple local patches of 64x64 pixels, and the size of the image is 512x512 pixels, the starting coordinate of the first local patch is (0, 0), the starting coordinate of the second local patch is (64, 0), and so on.
[0044] According to the cutting position of the target local tile, the position coordinates of the abnormal points are converted into absolute coordinates in the target alloy image, which can be obtained by adding the starting coordinates of the local tile and the relative coordinates of the abnormal points. For example, assuming that the starting coordinates of a local tile are (128, 128) and the relative coordinates of an abnormal point detected in the local tile are (32, 32), then the absolute coordinates of the abnormal point in the target alloy image are (128 + 32, 128 + 32) = (160, 160). In the target alloy image, according to the converted abnormal point coordinates, image processing techniques (such as drawing a rectangular frame according to the abnormal region composed of multiple abnormal points, or setting corresponding abnormal flag bits or abnormal labels at the abnormal points, etc.) are used to identify the defect region. For example, a rectangular frame can be drawn at the (160, 160) position of the target alloy image according to the distribution range of the abnormal points to identify the defect region at this position.
[0045] S140, according to the proportion of the defect region in the target alloy image, determining the surface processing defect degree of the target alloy after processing.
[0046] Specifically, for a rectangular defect region, the length and width can be directly calculated according to the coordinates of its contour, and then the area of the defect region is obtained. For a defect region with irregular shape, a fitting algorithm is needed to calculate its area. For example, a polygon fitting method can be used to approximate the irregular contour to one or more polygons, and then the area of the polygon is calculated according to the vertex coordinates. After calculating the area of the defect region, it is divided by the total rectangular area of the target alloy image (i.e. the length multiplied by the width of the image), to obtain the proportion of the defect region in the target alloy image. According to the size of the defect region proportion, the surface processing defect degree of the target alloy after processing is determined, and different proportion thresholds can be set to divide the defect degree levels, such as slight defect, moderate defect and severe defect, etc. For example, if the defect region proportion is less than 5%, it can be considered as a slight defect; the proportion between 5% and 20% is a moderate defect; and the proportion greater than 20% is a severe defect.
[0047] Alternatively, a pixel counting method can also be used to calculate the defect region proportion by counting the number of pixels in the abnormal defect feature map (regarding it as the area of the defect region) and the total number of pixels in the target alloy image, and then calculating the defect region proportion by (defect region pixel number / total pixel number of target alloy image) x 100%.
[0048] The technical scheme of the embodiment of the present application acquires the target alloy image after processing the target alloy, and according to the input data requirement of the pre-trained network model, the image is cut into a target local block and input into the model to obtain each abnormal defect feature map; according to the position coordinates of the abnormal points in each image and the cutting position of each target local block belonging to the target alloy image in the target alloy image, the defect area is marked; according to the proportion of the defect area in the target alloy image, the surface processing defect degree of the target alloy after processing is determined. The pre-trained network model is used to acquire the abnormal defect feature map of the target alloy image, and according to the proportion of the defect area in the target alloy image, the defect degree is quantified, the detection efficiency and precision of the alloy surface processing defect are improved, the alloy processing process can be monitored in real time, the labor cost is reduced, and the reliability of the alloy quality control is improved.
[0049] Optionally, on the basis of each of the above embodiments, the target alloy image corresponding to the processed target alloy can include:
[0050] According to the current environmental light condition and the surface reflection characteristic of the target alloy, an illumination light source is selected;
[0051] The target alloy placed at the shooting position is irradiated using the illumination light source, wherein an image acquisition device is fixedly arranged above the shooting position;
[0052] According to the size and texture of the target alloy, the image resolution and the acquisition frequency are selected, and the image acquisition device is set with image acquisition parameters according to the image resolution and the acquisition frequency;
[0053] The target alloy image corresponding to the target alloy is acquired using the image acquisition device with the completed parameter setting.
[0054] In the embodiments of the present application, the environmental lighting conditions can be specifically understood as the intensity, direction and color temperature of natural light or artificial light in the shooting site. Different environmental lighting conditions will affect the lighting effect on the surface of the target alloy, and further affect the quality of image acquisition. For example, in a relatively dark environment, a high-brightness lighting source can be selected; while in a strong outdoor environment, a low-brightness lighting source can be selected to avoid overexposure. The surface reflection characteristics of the target alloy can be specifically understood as the properties that affect the reflection effect of the lighting source on the surface of the alloy, including reflectivity and reflection spectrum, etc. For example, if the alloy surface is highly reflective, a non-direct soft light source can be used to reduce glare and reflection. If specific details of the alloy surface need to be captured, a directional light source or structured light can be used to highlight the texture and defects of the surface. The image resolution refers to the sharpness and detail performance of the image. For alloys with fine surface texture, small texture size or complex texture, a higher image resolution needs to be selected to capture more detailed information. For example, for the anodized texture on the surface of an aluminum alloy, a higher resolution can be set to clearly show the shape and distribution of the texture. The acquisition frequency refers to the speed of image acquisition. For scenes that require rapid detection or real-time monitoring, a higher acquisition frequency should be selected. For example, when detecting the surface processing status of an alloy workpiece on an automatic processing line in real time, the acquisition frequency needs to be set according to the processing speed to ensure that images of the workpiece surface at each processing stage can be continuously and timely acquired.
[0055] Specifically, according to the current environmental lighting conditions and the surface reflection characteristics of the target alloy, the illumination angle and distance of the light source also need to be adjusted according to the shape, size and surface characteristics of the target alloy to obtain the best lighting effect. The target alloy to be detected is placed at a pre-set shooting position. This position should ensure that the target alloy can be fully illuminated by the lighting source and be within the shooting range of the image acquisition device. The height and angle of the image acquisition device should be pre-adjusted according to the size of the target alloy and the shooting requirements. According to the size and texture of the target alloy, the image resolution and acquisition frequency are selected, and the image acquisition device is set with the parameters of image resolution and acquisition frequency. It can also include setting the parameters of camera exposure time, gain and white balance, etc. to ensure that high-quality images can be obtained at the selected resolution and frequency. After completing the parameter setting of the image acquisition device, the device is started for image acquisition. The device will shoot the target alloy according to the set parameters to obtain the target alloy image of the target alloy.
[0056] By adjusting the illumination light source and image acquisition parameters, different shooting environments and alloy materials can be adapted to meet diversified detection needs, enhance the adaptability and flexibility of detection, thereby improving the image acquisition quality and the precision of collecting detailed information of the alloy surface, meeting the needs of rapid detection or real-time monitoring, and improving the detection accuracy. The automated image acquisition process reduces manual intervention, reduces the complexity and workload of manual adjustment, improves production and processing efficiency, and accurate defect detection helps to discover and handle quality problems in a timely manner, reduces production costs and improves product quality.
[0057] Embodiment Two
[0058] Figure 2 The flowchart of another alloy surface processing defect detection method provided by Embodiment Two of the present application is a refinement of the operation of cutting the target alloy image into at least one target local block according to the input data requirements of the pre-trained network model in the above-mentioned embodiments. Specifically, it comprises: preprocessing the collected target alloy image, and cutting the processed image data into at least one target local block according to the input data requirements of the pre-trained network model.
[0059] Correspondingly, as shown in Figure 2 The method comprises:
[0060] S210, collecting a target alloy image corresponding to the processed target alloy.
[0061] S220, preprocessing the collected target alloy image, and cutting the processed image data into at least one target local block according to the input data requirements of the pre-trained network model.
[0062] Specifically, preprocessing the collected target alloy image can include: using a filter to remove noise in the image, such as Gaussian filtering or median filtering, etc., to reduce the influence of noise on subsequent processing; correcting the illumination of the image by histogram equalization or other methods, so that the brightness and contrast of the image are more uniform; scaling the image data to a unified numerical range (such as 0 to 1) by normalization, so as to eliminate the differences in illumination and contrast between different images, and improve the generalization ability; highlighting the defect features in the image by contrast enhancement (such as histogram equalization, etc.) or sharpening (such as Laplacian sharpening, etc.) or other methods, so that they are more easily recognized by the network model; and if necessary, the image can be converted from an RGB (Red Green Blue) image to other color images, such as grayscale images, etc., to facilitate subsequent feature extraction and analysis. By preprocessing the target alloy image, the quality of the target alloy image is improved, and the defect features in the image are highlighted, so that it is more in line with the input data requirements of the pre-trained network model.
[0063] S230, input each target local patch into the pre-trained network model to obtain an abnormal defect feature map corresponding to each target local patch.
[0064] The pre-trained network model adopts a semantic segmentation network based on segmentation of all.
[0065] S240, according to the position coordinates of the abnormal points in each abnormal defect feature map and the segmentation position of each target local patch belonging to the target alloy image in the target alloy image, mark a defect region in the target alloy image.
[0066] S250, according to the proportion of the defect region in the target alloy image, determine the surface processing defect degree of the target alloy after processing.
[0067] The technical scheme of the embodiment of the application, by collecting the target alloy image corresponding to the target alloy after processing, pre-processing the target alloy image, and according to the input data requirements of the pre-trained network model, the target alloy image is segmented into at least one target local patch; input each target local patch into the pre-trained network model to obtain an abnormal defect feature map corresponding to each target local patch; wherein the pre-trained network model adopts a semantic segmentation network based on segmentation of all; according to the position coordinates of the abnormal points in each abnormal defect feature map and the segmentation position of each target local patch belonging to the target alloy image in the target alloy image, mark a defect region in the target alloy image; according to the proportion of the defect region in the target alloy image, determine the surface processing defect degree of the target alloy after processing. Through image preprocessing, the image quality can be significantly improved, and the image features can be enhanced, so as to speed up the model inference speed, improve the detection accuracy of the model, obtain the abnormal defect feature map of the target alloy image through the pre-trained network model, and quantify the defect degree according to the proportion of the defect region in the target alloy image, improve the detection efficiency and precision of the alloy surface processing defect, and can monitor the alloy processing process in real time, reduce the labor cost, and improve the reliability of the alloy quality control.
[0068] Optionally, on the basis of each of the above embodiments, inputting the target local patch into the pre-trained network model to obtain an abnormal defect feature map corresponding to the target local patch can include:
[0069] input the current target local patch into the pre-trained network model;
[0070] The network model comprises: an image encoder and a decoder; the image encoder comprises a plurality of encoding layers, and each encoding layer comprises, in sequence, a linear layer, a position encoding layer, a multi-head self-attention mechanism layer, a feedforward neural network layer and a residual network layer; the decoder comprises a lightweight decoding layer, and the lightweight decoding layer comprises, in sequence, two transformer layers and two convolutional layers.
[0071] The image encoder in the network model is used for performing convolutional processing on the input current processing target local patch to obtain a plurality of two-dimensional feature maps.
[0072] The pixel values of the plurality of two-dimensional feature maps are rearranged into a one-dimensional vector, and the one-dimensional vector is mapped to a high-dimensional space by the image encoder to obtain an embedding vector.
[0073] The position encoding corresponding to the target local patch is added to the embedding vector, and the image encoder is used to obtain a target feature map.
[0074] The decoder in the network model is used for decoding the target feature map to obtain an abnormal defect feature map corresponding to the current processing target local patch.
[0075] In the embodiment of the application, the image encoder refers to the part of the network model used for extracting the features of the input image, which processes the input image through a series of encoding layers to convert the image data into feature representation. In the image encoder, the input current processing target local patch will be subjected to convolutional processing of a plurality of encoding layers to obtain a plurality of two-dimensional feature maps. These feature maps contain local feature information of the image, such as edges and textures.
[0076] The decoder refers to the part of the network model used for converting the feature map output by the encoder back to the image space, which processes the feature map through a series of decoding layers to generate an output image corresponding to the input image. In the decoder, the target feature map will be decoded to obtain an abnormal defect feature map corresponding to the current processing target local patch, and the feature map can reflect the abnormal defect region in the target local patch. The lightweight decoding layer refers to a decoding layer with a relatively simple structure, which decodes the feature map through less computing resources, and can efficiently decode the target feature map to obtain the abnormal defect feature map while maintaining the computing efficiency of the model.
[0077] Transformer layer is a neural network layer based on the Transformer architecture, which uses self-attention mechanisms and feedforward neural networks to process and extract features from input data. In the lightweight decoding layer, the Transformer layer can further fuse and transform the feature maps, enhancing the model's understanding and expression of image features. Convolutional layer is a neural network layer that extracts and processes features from input data through convolution operations. In the lightweight decoding layer, the convolutional layer can extract local features and process spatial information from the feature maps to generate the final abnormal defect feature map.
[0078] Linear layer refers to the fully connected layer in the network model, which maps input data to output data through linear transformation, i.e., linear combination of input feature data, providing a basis for subsequent nonlinear transformation and feature extraction. Position encoding layer can be understood as follows: a part of the network model that adds position information to the input data, enabling the model to recognize the position relationship of the data in the sequence. In the image encoder, the position encoding layer adds the position encoding corresponding to the target local patch to the embedding vector, enabling the model to understand the position of the image patch in the original image, thus better performing feature extraction and analysis.
[0079] Multi-head self-attention mechanism layer refers to a neural network layer in the network model based on the self-attention mechanism, which divides the input data into multiple parts and independently calculates the self-attention in each part, then combines the results. This layer enables the model to simultaneously focus on different regions and features of the image when processing image features, enhancing the model's understanding of global information and feature fusion capabilities. Feedforward neural network layer refers to a neural network layer in the network model composed of multiple neurons, which transmits and processes information in a feedforward manner. In the encoding layer, the feedforward neural network layer can perform nonlinear transformation and feature extraction on the input feature data, further enhancing the expressiveness of the features. Residual network layer refers to the part of the network model constructed by introducing residual connections (i.e., directly adding input data and output data), which can alleviate the gradient vanishing problem in neural networks, enabling the model to perform deeper feature extraction and learning.
[0080] Specifically, the image encoder in the network model extracts local feature information through convolution processing, obtains multiple two-dimensional feature maps, flattens the pixel values of the multiple two-dimensional feature maps, and rearranges them into a one-dimensional vector. For example, for a feature map with a shape of (channel number, height, width), it is flattened into a one-dimensional vector with a shape of (channel number x height x width). Through the linear layer (such as the fully connected layer) in the image encoder, the one-dimensional feature vector is mapped to a high-dimensional space to obtain an embedding vector. The embedding vector is a high-dimensional representation of the image features, which can capture deeper feature information of the image. The position encoding corresponding to the target local patch contains the position information of the image patch in the original image, which can be added to the embedding vector through an addition operation to integrate the position information into the feature representation. The embedding vector processed by the image encoder obtains a target feature map containing semantic information, position information, and local feature information of the image. In the lightweight decoding layer of the decoder, the Transformer layer performs feature fusion and transformation on the feature map to enhance the feature expression capability. The convolutional layer performs local feature extraction and spatial information processing on the transformed feature map. Finally, an abnormal defect feature map corresponding to the current processing target local map is generated, and the pixel value distribution of the feature map reflects the position and range of the defect area, providing a basis for subsequent defect detection and analysis.
[0081] By inputting the target local patch into the pre-trained network model, the image processing is automated without human intervention, which can quickly process a large amount of image data, improve the efficiency of image processing, save labor costs and time costs, and is conducive to real-time monitoring of the alloy processing process. At the same time, the model efficiently extracts feature information from the image through the cooperation of the image encoder and the decoder. The encoder quickly captures local and global features of the image using structures such as convolutional layers and multi-head self-attention mechanism layers, and the decoder efficiently decodes the feature map into a target image through the lightweight decoding layer. The entire process has relatively small computational complexity and fast processing speed. In addition, the multi-head self-attention mechanism layer realizes the fusion of multi-dimensional features, enabling the model to more comprehensively understand the image content, capture complex features and subtle changes, and enhance the expression capability of image features. By introducing position encoding, the pixel value distribution of the abnormal defect feature map can clearly reflect the position and range of the defect, making the defect detection more accurate, reducing the false detection and missed detection, improving the accuracy of defect detection, and improving the reliability of alloy quality control.
[0082] Typically, the image encoder can include twelve encoding layers, and the decoder is a lightweight decoding layer. In a specific example, a full-size alloy surface RGB image can be processed into multiple local RGB images of 1024*1024*3; the processed local RGB images are convolved by 16*16 to obtain 768 feature maps with a shape of 64*64; each 64*64 image block is flattened into a 768-dimensional vector and mapped to a higher-dimensional space through a linear layer to obtain an embedding vector; in order to introduce the position information of the image block, position encoding is added to each embedding vector, wherein the position encoding vector has the same dimension as the embedding vector and is used to represent the information of each position in the input sequence; the embedding vector and the position encoding information are input into the Transformer encoding layer, in each encoder layer, the embedding vector is processed through multi-head self-attention mechanism and feedforward neural network to capture the relationship between image blocks and feature representation, and after twelve encoding layers, different levels of information and features of the image can be captured; the feature map obtained by the twelfth encoding layer is further mapped and reduced in dimension through two convolution layers to obtain a feature map of (1, 256, 64, 64), wherein (1, 256, 64, 64) represents (batch size, channel number, height, width), that is, the current processing is an image, the feature map has 256 channels, and the height and width of the feature map are both 64 pixels; the (1, 256, 64, 64) feature map obtained by the encoder is input into the decoder, after capturing the global and local relationships in the feature map through two transformer layers, the feature map is further processed through two convolution layers to achieve the effect of doubling the length and width and halving the channel number, and finally the segmented abnormal defect feature map is output.
[0083] Further, on the basis of each of the above embodiments, before each target local block is input into the pre-trained network model, the following steps can be further included:
[0084] According to the defect detection requirement, the loss function in the pre-trained network model is adjusted, the labeled data related to each target local block is input into the pre-trained network model, and the model parameters are adjusted.
[0085] Generally, different defect types and detection scenarios may require the model to focus on different performance indicators. For example, in some defect type detection scenarios, more attention needs to be paid to the accurate identification of defect boundaries by the model, while in some cases, more attention needs to be paid to the overall segmentation effect of the defect area by the model. By adjusting the loss function, the training target of the model can be matched with the actual defect detection requirements, thereby improving the performance of the model on specific tasks. Typically, the loss function can take the form of a hybrid loss function that combines multiple loss functions to balance the performance requirements in different aspects. Preferably, a hybrid loss function that combines cross-entropy loss and Dice Loss (segmentation loss) can be used. Cross-entropy loss measures the difference between predicted probabilities and true labels in classification tasks by calculating the distance between predicted probability distribution and true distribution, which is suitable for classification tasks and helps to improve the classification accuracy of the model for defect types. Dice Loss is used to measure the overlap between the predicted result and the true label, which is suitable for handling unbalanced data sets and helps to improve the segmentation accuracy of the model for defect areas. By combining these two losses, the model can optimize both classification accuracy and segmentation effect, thereby improving overall performance. Accordingly, the calculation method of the hybrid loss function can be: calculating the cross-entropy loss between the predicted probability distribution of the model for each target local patch and the true label; calculating the Dice coefficient between the model's predicted segmentation mask and the true segmentation mask, and then converting it to Dice Loss, where the higher the Dice coefficient, the higher the overlap between the predicted segmentation area and the true segmentation area, and the lower the Dice Loss; and weighted sum of cross-entropy loss and Dice Loss to obtain the final hybrid loss function. The weights can be adjusted according to the actual requirements of defect types and detection scenarios to balance the importance of classification accuracy and segmentation effect.
[0086] In the embodiments of the present application, the labeled data refers to training samples with true labels. The labeled data related to each target local patch is input into the pre-trained network model, and the model can calculate the loss between the predicted result and the true label, adjust its parameters using the gradient descent algorithm, and make the predicted result gradually approach the true label, thereby improving the performance of the model. Accordingly, the network can also include an adapter for injecting domain knowledge into the large model.
[0087] Specifically, the annotation data related to each target local patch is input into the pre-trained network model, and the annotation data includes images and corresponding defect labels (such as defect types, defect region segmentation masks, etc.). The model calculates the loss between the predicted results and the true labels according to the hybrid loss function, and uses the gradient descent algorithm to adjust the parameters of the model according to the gradient information of the loss function, so that the value of the loss function gradually decreases until it meets the preset requirements. A small amount of labeled data is used to fine-tune the large-scale pre-trained segmentation network, so that the training target of the model is accurately matched with the actual task requirements, and the pertinence and adaptability of the model are enhanced. The use of the hybrid loss function enhances the classification and segmentation performance of the model, improves the recognition accuracy of the defect type and the segmentation precision of the defect region, and makes the prediction results more accurate and reliable. Through the form adjustment of the hybrid loss function and the optimization training of the labeled data, the generalization ability and robustness of the model are significantly improved, so that it can better cope with new and unlabeled image data, complex environmental conditions and data quality changes, improve the detection efficiency and precision of alloy surface processing defects, and can monitor the alloy processing process in real time, reduce the labor cost, and improve the reliability of the alloy quality control.
[0088] Optionally, based on the above embodiments, according to the position coordinates of the abnormal points in each abnormal defect feature map and the segmentation position of each target local patch in the target alloy image to which the abnormal defect feature map belongs, the defect region in the target alloy image is marked, which can include:
[0089] The abnormal defect feature maps are spliced according to the position encoding, the spliced abnormal defect feature maps are processed through morphological operations, and the position coordinates of the abnormal points in each abnormal defect feature map and the segmentation position of each target local patch in the target alloy image to which the abnormal defect feature map belongs are determined according to the position encoding, and the defect region in the target alloy image is marked.
[0090] Specifically, each abnormal defect feature map is stitched according to its position encoding in the original image. The position encoding records the specific position of each feature map in the target alloy image. By stitching, the scattered feature maps can be integrated into a complete image, thereby forming a continuous defect area view in the target alloy image. Then, morphological operations such as erosion and dilation can be used to optimize the feature maps. The erosion operation can remove small noise points, while the dilation operation can fill the holes in the defect area, making the defect boundary more clear. Combined with the position encoding, the position of the abnormal point in each abnormal defect feature map and the position of the target local block to which the feature map belongs in the target alloy image can be determined. In the target alloy image, according to the converted abnormal point coordinates and the segmentation position of the target local block, image processing techniques such as drawing a rectangular frame or marking a point are used to identify the defect area. Through position encoding stitching and morphological operation processing, the defect area in the target alloy image can be more accurately identified, reducing false positives and missed detections, thereby improving the accuracy of defect detection. At the same time, the defect area is clearly identified in the target alloy image, making the defect information easy to see, facilitating subsequent defect analysis and defect degree judgment, and enhancing the readability of the image.
[0091] Specific application scenarios
[0092] Alloys are widely used in industries, construction, aerospace and other fields. Among them, aluminum alloy is a widely used lightweight and corrosion-resistant metal material. In order to improve the corrosion resistance, hardness and wear resistance of aluminum alloy, anodic oxidation treatment is often required. Anodic oxidation is a method of generating an oxide layer on the surface of aluminum alloy through an electrolytic process, which can enhance the surface hardness and corrosion resistance of aluminum alloy. However, during the anodic oxidation process, some defects may occur, such as uneven oxide layer thickness, holes, cracks and poor sealing. These defects not only affect the appearance quality of aluminum alloy products, but also may reduce their performance and service life. Therefore, timely detection and evaluation of anodic oxidation layer defects are crucial for ensuring product quality.
[0093] In the production process of aluminum and aluminum alloy oxidation, various defects generated can be mainly divided into surface defects of oxidized surface treatment products, shape, size and appearance performance defects of oxidized surface treatment products, etc. Traditional aluminum alloy surface anodic oxidation defect detection needs professional personnel to measure on site. Due to the relatively limited detection accuracy of the human eye, it is difficult to find small, hidden or deep defects, and there are problems of high labor cost and limited detection accuracy. These defects can have a serious impact on the performance and safety of the product. In addition, manual detection is usually intermittent, making it difficult to achieve real-time monitoring of the production process, which can lead to missed defects in the production process, increasing the risk of product quality and reducing the reliability of alloy quality control.
[0094] With the rapid development of deep learning, using a deep learning-based method to replace manual detection has become one of the trends in alloy surface processing defect detection. The current mainstream method is to collect two-dimensional picture data and detect alloy surface defects by using a neural network-based semantic segmentation method. However, in this defect detection field, the number of defect samples is relatively small, which makes it very challenging to build an accurate and robust deep learning model. Secondly, it is also quite difficult to accurately segment and label the defect samples. Since defects often have complex shapes and textures, manual accurate labeling requires a lot of time and effort, and is easily affected by subjective factors, leading to inconsistency and inaccuracy in labeling, which reduces the reliability of alloy quality control.
[0095] To solve the above problems, an alloy surface processing defect detection method is provided in an embodiment of the present application, Figure 3 is a flowchart of an alloy surface processing defect detection and evaluation method suitable for an embodiment of the present application. Taking the aluminum alloy surface anodic oxidation defect detection and defect evaluation application scenario as an example, the method can specifically include:
[0096] 1. Obtain the RGB image of the target aluminum alloy plate surface
[0097] According to the environmental lighting conditions and surface reflection characteristics, a suitable light source is selected to illuminate the aluminum alloy surface to ensure the quality of the image. An image acquisition device is fixedly arranged above the shooting position of the aluminum alloy. During the acquisition process, appropriate image resolution and acquisition frequency are selected according to the size and texture of the aluminum alloy surface to ensure that the image details of each region are clear and comprehensive. The range of the target aluminum alloy plate is determined, and the RGB image of the target aluminum alloy plate surface is obtained by the acquisition device. The range of the target aluminum alloy plate can be described by a rectangle uniquely determined by four vertices.
[0098] 2. Obtain the abnormal defect feature map
[0099] The collected aluminum alloy surface image is preprocessed, such as denoising, normalization and contrast enhancement, to ensure good image quality and facilitate subsequent analysis. The processed full-size aluminum alloy surface RGB image is divided into local RGB images to ensure that each image size is suitable for input into the segmentation network for processing. The local RGB image is input into the semantic segmentation network to obtain the abnormal defect feature map of the segmentation result. The semantic segmentation network can use the SAM semantic segmentation network of the large-scale pre-training model Segment Anything, Figure 4 is a network structure diagram of a large-scale pre-training adaptive network suitable for embodiments of the present application. The network includes an Adapter module for injecting domain knowledge into a large model, and an encoder and a decoder for extracting 2D image feature information to generate a corresponding abnormal defect feature map. The input of the network is the processed local RGB image, and the output of the network is the abnormal defect feature map of the segmentation result.
[0100] The image encoder includes a plurality of encoding layers, each encoding layer including a linear layer, a position encoding layer, a multi-head self-attention mechanism, a feedforward neural network layer and a residual connection. Each encoding layer can gradually extract and convert input features to capture different levels of information and features of the image. The decoder includes a lightweight decoding layer, and the decoding layer includes two transformer layers and two convolution layers. After each convolution layer, the feature map is doubled in length and width and halved in channel number. Optionally, the image encoder can include twelve encoding layers, and the decoder can be a lightweight decoding layer.
[0101] In addition, the obtained local RGB image block can be combined with a small amount of labeled data to fine-tune the large-scale pre-trained segmentation network, so that it can recognize and segment the anodic film defects on the surface of the aluminum alloy. During the fine-tuning process, the pre-trained model is injected with domain-specific knowledge, the loss function is adjusted according to the characteristics of the aluminum alloy surface defects, and the diversity of the enhanced data set is increased to improve the adaptability of the model in actual production. The fine-tuned segmentation network can infer new local RGB image blocks and output pixel-level abnormal defect feature maps to extract key information such as spatial distribution, shape and size of defects in subsequent steps.
[0102] Preferably, the full-size aluminum alloy plate surface RGB image can be processed into a 1024*1024*3 local RGB image. Correspondingly, obtaining the abnormal defect feature map can include:
[0103] (1) The obtained full-size aluminum alloy surface RGB image is processed into a plurality of 1024*1024*3 local RGB images.
[0104] (2) The processed local RGB image is convolved by 16*16 to obtain 768 feature maps with a shape of 64*64.
[0105] (3) Each 64x64 image is rearranged into a 768-dimensional vector and mapped to a higher-dimensional space through a linear layer to obtain an embedding vector.
[0106] (4) Since the ViT (Vision Transformer) model has no ability to process pixel position information in the image, in order to introduce the position information of the image block, a position encoding is added to each embedding vector. The position encoding is usually a vector with the same dimension as the embedding vector, which represents the information of each position in the input sequence.
[0107] (5) The embedding vector and the position encoding are input into the Transformer encoder layer. In each encoder layer, the embedding vector is processed through the multi-head self-attention mechanism and the feedforward neural network, so as to capture the relationship between the image blocks and the feature representation. After passing through twelve encoder layers, the different levels of information and features of the image can be captured.
[0108] (6) The feature map obtained by the twelfth encoder layer is further mapped and reduced in dimension through two convolutional layers to obtain a feature map of (1, 256, 64, 64).
[0109] (7) The (1, 256, 64, 64) feature map obtained by the encoder is put into the decoder. After capturing the global and local relationships in the feature map through two transformer layers, the feature map is doubled in length and width and halved in channel number through two convolutional layers, and finally the segmented mask map (equivalent to the abnormal defect feature map in the previous text) is output.
[0110] 3. Evaluate the defect state
[0111] According to the output abnormal defect feature map, the defect area is finely processed through morphological operations such as erosion and dilation. The erosion operation helps to remove unnecessary small noise in the image, and the dilation operation helps to fill the gaps in the defect area, so as to ensure the integrity of the defect boundary. Through these operations, the edge of the defect area can be accurately identified, and errors caused by image noise or illumination changes can be avoided. After identifying the edge of the defect, the length and width of the defect map are calculated to obtain the area of the defect area. According to the area of the defect map and the total area of the aluminum alloy surface, the proportion of the defect on the overall surface is calculated. The higher the proportion, the more serious the defect.
[0112] The alloy surface processing defect detection method provided by the embodiment of the present application uses a large-scale pre-training model, fully utilizes large-scale prior knowledge, and overcomes the problems of lack of deep understanding of complex defects and easy overfitting and poor generalization ability of traditional detection methods in a small sample case; by introducing an adapter to inject domain professional knowledge into the large-scale pre-training model, the model is adapted in the field of alloy surface processing defect detection, such as the field of aluminum alloy surface anodic oxidation defect detection, the semantic understanding and feature extraction capability of the model for alloy surface processing defects are improved, and the model can accurately locate and segment defects of various shapes, sizes and textures; through the combination of the large-scale pre-training model and the adapter, after adaptation with a small amount of data, the model has strong generalization capability, so that the model can also accurately label defects on unknown data, the degree of defects is quantified by calculating the proportion, the detection efficiency and precision of the alloy surface processing defects are improved, the alloy processing process can be monitored in real time, the labor cost is reduced, and the reliability of the alloy quality control is improved.
[0113] Embodiment three
[0114] Figure 5 A structural schematic diagram of an alloy surface processing defect detection device provided by the embodiment three of the present application is shown in the figure. Figure 5 As shown in the figure, the device comprises an image segmentation module 510, a feature map acquisition module 520, a defect identification module 530 and a degree determination module 540, wherein:
[0115] The image segmentation module 510 is configured to collect a target alloy image corresponding to a processed target alloy, and segment the target alloy image into at least one target local block according to the input data requirements of the pre-trained network model;
[0116] The feature map acquisition module 520 is configured to input each target local block into the pre-trained network model to obtain an abnormal defect feature map corresponding to each target local block; wherein the pre-trained network model adopts a semantic segmentation network based on segmentation of everything;
[0117] The defect identification module 530 is configured to identify a defect region in the target alloy image according to the position coordinates of the abnormal points in each abnormal defect feature map and the segmentation position of each abnormal defect feature map in the target alloy image;
[0118] The degree determination module 540 is configured to determine the surface processing defect degree of the target alloy after processing according to the proportion of the defect region in the target alloy image.
[0119] The technical scheme of the embodiment of the present application acquires the target alloy image after processing of the target alloy, and according to the input data requirement of the pre-trained network model, cuts the image into a target local block and inputs into the model to obtain each abnormal defect feature map; according to the position coordinates of the abnormal points in each image and the cutting position of each image in the target alloy image, the defect area is marked; according to the proportion of the defect area in the target alloy image, the surface processing defect degree of the target alloy after processing is determined. The pre-trained network model is used to acquire the abnormal defect feature map of the target alloy image, and according to the proportion of the defect area in the target alloy image, the defect degree is quantified, the detection efficiency and precision of the alloy surface processing defect are improved, the alloy processing process can be monitored in real time, the labor cost is reduced, and the reliability of the alloy quality control is improved.
[0120] On the basis of the above embodiments, the image cutting module 510 is specifically used for:
[0121] According to the current environmental light condition and the surface reflection characteristic of the target alloy, an illumination light source is selected;
[0122] The target alloy placed at the shooting position is irradiated using the illumination light source, wherein an image acquisition device is fixedly arranged above the shooting position;
[0123] According to the size and texture of the target alloy, the image resolution and the acquisition frequency are selected, and the image acquisition device is set with image acquisition parameters according to the image resolution and the acquisition frequency;
[0124] The image acquisition device with the completed parameter setting is used to acquire the target alloy image corresponding to the target alloy.
[0125] Further, on the basis of the above embodiments, the image cutting module 510 can be specifically used for:
[0126] The acquired target alloy image is preprocessed, and the preprocessed image data is cut into at least one target local block according to the input data requirement of the pre-trained network model.
[0127] On the basis of the above embodiments, the feature map acquisition module 520 is specifically used for:
[0128] The current processed target local block is input into the pre-trained network model;
[0129] The network model comprises an image encoder and a decoder; the image encoder comprises a plurality of encoding layers, and each encoding layer comprises, in sequence, a linear layer, a position encoding layer, a multi-head self-attention mechanism layer, a feedforward neural network layer and a residual network layer; the decoder comprises a lightweight decoding layer, and the lightweight decoding layer comprises, in sequence, two transformer layers and two convolutional layers.
[0130] The image encoder in the network model is used for performing convolutional processing on the input current processing target local patch to obtain a plurality of two-dimensional feature maps.
[0131] The pixel values of the plurality of two-dimensional feature maps are rearranged into a one-dimensional vector, and the one-dimensional vector is mapped to a high-dimensional space by the image encoder to obtain an embedding vector.
[0132] The position encoding corresponding to the target local patch is added to the embedding vector, and the image encoder is used to obtain a target feature map.
[0133] The decoder in the network model is used for decoding the target feature map to obtain an abnormal defect feature map corresponding to the current processing target local patch.
[0134] Further, on the basis of each of the above embodiments, the alloy surface processing defect detection device can further comprise:
[0135] A parameter adjustment module is configured to, before inputting each target local patch into the pre-trained network model, adjust a loss function in the pre-trained network model according to a defect detection requirement, input labeled data related to each target local patch into the pre-trained network model, and adjust model parameters.
[0136] On the basis of each of the above embodiments, the defect identification module 530 is specifically configured to:
[0137] The abnormal defect feature maps are spliced according to the position encoding, the spliced abnormal defect feature maps are processed through morphological operations, and the position coordinates of the abnormal points in each abnormal defect feature map and the split positions of each abnormal defect feature map in the target alloy image are determined in combination with the position encoding, so as to identify a defect region in the target alloy image.
[0138] The alloy surface processing defect detection device provided in the embodiments of the present application can execute the alloy surface processing defect detection method provided in any of the embodiments of the present application, and has the corresponding function modules and beneficial effects of the execution method.
[0139] In the technical solutions of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information comply with relevant laws and regulations and do not violate public order and good customs.
[0140] Example 4
[0141] Figure 6 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0142] like Figure 6 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0143] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0144] The processor 11 may be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the method for detecting alloy surface processing defects, namely:
[0145] collecting a target alloy image corresponding to the target alloy after the target alloy is collected and processed, and cutting the target alloy image into at least one target local block according to input data requirements of the pre-trained network model;
[0146] inputting each target local block into the pre-trained network model to obtain an abnormal defect feature map corresponding to each target local block; wherein the pre-trained network model adopts a semantic segmentation network based on segmentation of all things;
[0147] identifying a defect region in the target alloy image according to position coordinates of abnormal points in each abnormal defect feature map and cutting positions of each target local block belonging to the abnormal defect feature map in the target alloy image;
[0148] determining a surface processing defect degree of the target alloy after processing according to a proportion of the defect region in the target alloy image.
[0149] In some embodiments, the method of detecting surface processing defects of an alloy can be implemented as a computer program tangibly embodied in a computer readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded onto RAM 13 and executed by processor 11, one or more steps of the method of detecting surface processing defects of an alloy described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the method of detecting surface processing defects of an alloy by any other suitable means, such as by means of firmware.
[0150] The various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (PLD), a computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0151] Computer programs for implementing the methods of the present application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer program, when executed, enables the functions / acts specified in the flowcharts and / or block diagrams to be implemented. The computer program can be executed entirely on a machine, partially on a machine, partially on a machine as a standalone software package and partially on a remote machine or entirely on a remote machine or server.
[0152] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. A computer-readable storage medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of a machine-readable storage medium will include one or more lines of electrical connections, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0153] To provide for interaction with a user, the systems and techniques described here can be implemented on an electronic device having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0154] The systems and techniques described herein can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described herein, or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0155] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. A server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service.
[0156] It should be understood that the various forms of flow shown above can be re-ordered, added to, or deleted from without departing from the scope of the present disclosure. For example, the steps recited in the present disclosure can be performed in parallel, in series, or in a different order, and the present disclosure is not limited herein as long as the desired results of the technical solutions of the present disclosure can be achieved.
[0157] The above detailed description does not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A method for detecting alloy surface processing defects, characterized in that: include: Acquire a target alloy image corresponding to the processed target alloy, and segment the target alloy image into at least one target local image block according to the input data requirements of the pre-trained network model; Input each target local block into the pre-trained network model to obtain the abnormal defect feature map corresponding to each target local block; wherein the pre-trained network model adopts a semantic segmentation network based on segmentation of everything; Identify defect areas in the target alloy image according to the position coordinates of the abnormal points in each abnormal defect feature map and the segmentation positions of the target local image blocks to which each abnormal defect feature map belongs in the target alloy image; According to the proportion of the defect area in the target alloy image, the degree of surface processing defects of the target alloy after processing is determined.
2. The method according to claim 1, characterized in that Acquire target alloy images corresponding to the processed target alloy, including: Select the lighting source based on the current ambient lighting conditions and the surface reflectance characteristics of the target alloy; Using an illumination light source to irradiate a target alloy placed at a shooting position, wherein an image acquisition device is pre-fixed above the shooting position; Selecting an image resolution and an acquisition frequency according to the size and texture of the target alloy, and setting image acquisition parameters of an image acquisition device according to the image resolution and the acquisition frequency; Use the image acquisition device with completed parameter settings to acquire a target alloy image corresponding to the target alloy.
3. The method according to claim 1, characterized in that According to the input data requirements of the pre-trained network model, the target alloy image is divided into at least one target local block, including: The collected target alloy image is preprocessed, and the processed image data is divided into at least one target local block according to the input data requirements of the pre-trained network model.
4. The method according to claim 1, wherein Input the target local image block into the pre-trained network model to obtain the abnormal defect feature map corresponding to the target local image block, including: Input the currently processed target local patch into the pre-trained network model; The network model includes: an image encoder and a decoder; the image encoder includes multiple encoding layers, each encoding layer includes a linear layer, a position encoding layer, a multi-head self-attention mechanism layer, a feedforward neural network layer and a residual network layer connected in sequence; the decoder includes a lightweight decoding layer, which includes two transformer layers and two convolutional layers connected in sequence; Through the image encoder in the network model, the input current processing target local block is convolved to obtain multiple two-dimensional feature maps; Rearranging the pixel values of the multiple two-dimensional feature maps into a one-dimensional vector, and mapping it to a high-dimensional space through an image encoder to obtain an embedding vector; The position code corresponding to the target local patch is added to the embedding vector, and the target feature map is obtained through the image encoder; The target feature map is decoded by the decoder in the network model to obtain the abnormal defect feature map corresponding to the current processing target local map.
5. The method according to any one of claims 1 to 4, characterized in that Before inputting each target local patch into the pre-trained network model, it also includes: According to the defect detection requirements, the loss function in the pre-trained network model is adjusted, and the labeled data related to each target local block is input into the pre-trained network model to adjust the model parameters.
6. The method according to any one of claims 1 to 4, characterized in that According to the position coordinates of the abnormal points in each abnormal defect feature map and the segmentation position of the target local image block to which each abnormal defect feature map belongs in the target alloy image, the defect area is identified in the target alloy image, including: Each abnormal defect feature map is spliced according to the position coding, and the spliced abnormal defect feature map is processed through morphological operations. The position coordinates of the abnormal points in each abnormal defect feature map and the segmentation position of the target local block to which each abnormal defect feature map belongs in the target alloy image are determined in combination with the position coding, and the defect area is identified in the target alloy image.
7. A device for detecting alloy surface processing defects, characterized in that: include: An image segmentation module is used to collect a target alloy image corresponding to the processed target alloy and segment the target alloy image into at least one target local image block according to the input data requirements of the pre-trained network model; The feature map acquisition module is used to input each target local block into the pre-trained network model to obtain the abnormal defect feature map corresponding to each target local block; wherein the pre-trained network model adopts a semantic segmentation network based on segmentation of everything; A defect identification module is used to identify defect areas in the target alloy image based on the position coordinates of the abnormal points in each abnormal defect feature map and the segmentation position of the target local image block to which each abnormal defect feature map belongs in the target alloy image; The degree determination module is used to determine the degree of surface processing defects of the target alloy after processing based on the proportion of the defect area in the target alloy image.
8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor. The computer program is executed by the at least one processor so that the at least one processor can perform the method for detecting alloy surface processing defects according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the method for detecting alloy surface processing defects according to any one of claims 1 to 6 when executed.
10. A computer program product, characterized in that The computer program product comprises a computer program, which, when executed by a processor, implements the method for detecting alloy surface processing defects according to any one of claims 1 to 6.
Citation Information
Patent Citations
Textile fabric surface defect detection method based on artificial intelligence and Gaussian mixture model
CN114529538A
Flywheel panel semi-finished product surface defect detection method based on neural network
CN115018764A
Insulator defect detection method and device, medium and program product
CN115984226A
Metal plate cutting processing method and system based on image segmentation
CN116630255A
Rail surface defect detection method and device, electronic equipment and storage medium
CN117952946A