Die defect detection method and device based on deep learning
By applying a significant focus strategy in mold defect detection to perform feature fusion and element relationship classification, the problem of low precision in mold defect detection in the prior art is solved, and more efficient defect identification and positioning is achieved.
Patent Information
- Application Number
- CN202510204242.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-02-24
AI Technical Summary
The existing deep learning-based mold defect detection methods are difficult to accurately identify and locate defects when processing complex mold machine vision images, especially when there are many image elements and complex relationships, the defect detection accuracy is not high.
By acquiring candidate feature representations of image elements in mold machine vision images, feature fusion and element relationship classification are used to determine the involvement relationship between image elements, and optimize feature representation to improve detection accuracy.
It improves the accuracy and efficiency of mold defect detection, can more accurately identify and position defects on the mold, and improves production quality and efficiency.
Smart Images

Figure CN120259177A_ABST
Abstract
Description
Technical Field
[0001] This application relates to, but is not limited to, the fields of image processing and machine learning technologies, and particularly relates to a method and device for mold defect detection based on deep learning. Background Art
[0002] With the continuous progress of industrial manufacturing technology, molds play a crucial role in the production of various products. However, various defects may occur during the use of molds, and these defects will seriously affect the quality and production efficiency of products. Therefore, it is particularly important to accurately detect defects in molds.
[0003] Traditional mold defect detection methods mainly rely on manual inspection or simple image processing techniques. These methods are not only inefficient but also inaccurate and are easily affected by human factors and complex environments. In recent years, with the rapid development of deep learning technology, it has demonstrated excellent performance in image processing and vision tasks, providing a new solution idea for mold defect detection.
[0004] In the field of deep learning, neural network models can automatically learn and extract features in images, and then perform tasks such as image classification, recognition, and detection. However, existing deep learning-based mold defect detection methods often have difficulty accurately identifying and locating defects when dealing with complex mold machine vision images, especially when there are many image elements (such as when the mold has complex textures) and complex relationships.
[0005] In addition, the element relationships in mold machine vision images are crucial for defect detection. There may be complex involvement relationships between different image elements, and these relationships play a key role in accurately judging the type and location of defects. Therefore, how to improve the accuracy of deep learning-based mold defect detection methods, especially when dealing with complex mold machine vision images, has become an urgent problem to be solved. Summary of the Invention
[0006] In view of this, at least one embodiment of this application provides a method and device for mold defect detection based on deep learning. The technical solution of the embodiment of this application is implemented as follows: On the one hand, an embodiment of this application provides a method for mold defect detection based on deep learning, and the method includes: Obtain a mold machine vision image, the mold machine vision image is composed of a plurality of image regions, and the mold machine vision image includes X image elements, where X>1; Obtain a candidate feature representation of each of the X image elements, and one image element corresponds to at least one candidate feature representation; Perform a feature fusion operation on the candidate feature representations of each image element according to the saliency focusing strategy to obtain the adjusted candidate feature representation of each image element; Based on the adjusted candidate feature representation of each image element, determine the image element feature representation of each image element; Perform an element relationship classification on the image element feature representations of the X image elements to determine the element involvement relationship of the X image elements in the die machine vision image.
[0007] In some embodiments, any one of the X image elements is regarded as the first image element, the first image element corresponds to Y candidate feature representations, and any one of the Y candidate feature representations of the first image element is regarded as the u-th candidate feature representation of the first image element; both u and Y are natural numbers greater than 0, and at the same time u ≤ Y; The performing a feature fusion operation on the candidate feature representations of each image element according to the saliency focusing strategy to obtain the adjusted candidate feature representation of each image element includes: Obtain the commonality measurement coefficients between the u-th candidate feature representation of the first image element and each of the Y candidate feature representations of the first image element, and obtain the set of commonality measurement coefficients corresponding to the u-th candidate feature representation; Obtain a preset parameter, and perform a downsampling operation on each of the commonality measurement coefficients in the set of commonality measurement coefficients corresponding to the u-th candidate feature representation based on the preset parameter to obtain the downsampling operation results of each of the commonality measurement coefficients in the set of commonality measurement coefficients corresponding to the u-th candidate feature representation; Perform a normalization operation on the downsampling operation results of each of the commonality measurement coefficients in the set of commonality measurement coefficients corresponding to the u-th candidate feature representation to obtain the influence coefficient cluster corresponding to the u-th candidate feature representation; Perform an eccentric adjustment process on the Y candidate feature representations of the first image element according to the influence coefficient cluster corresponding to the u-th candidate feature representation to obtain the adjusted u-th candidate feature representation of the first image element.
[0008] In some embodiments, any one of the Y candidate feature representations of the first image element is regarded as the v-th candidate feature representation of the first image element; v is a natural number greater than 0, and at the same time v ≤ Y; the obtaining the commonality measurement coefficients between the u-th candidate feature representation of the first image element and each of the Y candidate feature representations of the first image element, and obtaining the set of commonality measurement coefficients corresponding to the u-th candidate feature representation includes: Obtain a first variable two-dimensional array and a first variable one-dimensional array; Based on the two-dimensional array of the first variable, an integration operation is performed on the u-th candidate feature representation and the v-th candidate feature representation of the first image element to obtain an integration operation result; Perform a non-linear transformation on the integration operation result, and determine a commonality measurement coefficient between the u-th candidate feature representation and the v-th candidate feature representation based on the non-linearly transformed integration operation result and the one-dimensional array of the first variable; Fuse the commonality measurement coefficient between the u-th candidate feature representation and the v-th candidate feature representation into the set of commonality measurement coefficients corresponding to the u-th candidate feature representation to obtain the set of commonality measurement coefficients corresponding to the u-th candidate feature representation; The influence coefficient cluster corresponding to the u-th candidate feature representation includes Y influence coefficients, and the Y influence coefficients correspond one-to-one to the Y candidate feature representations corresponding to the first image element; The eccentric adjustment process of the Y candidate feature representations of the first image element according to the influence coefficient cluster corresponding to the u-th candidate feature representation to obtain the u-th candidate feature representation of the adjusted first image element includes: Multiply the Y candidate feature representations corresponding to the first image element by the Y influence coefficients corresponding to the u-th candidate feature representation to obtain an influence adjustment result of the Y candidate feature representations corresponding to the first image element; Sum the influence adjustment results of the Y candidate feature representations corresponding to the first image element to obtain the u-th candidate feature representation of the adjusted first image element.
[0009] In some embodiments, the obtaining of the candidate feature representation of each of the X image elements includes: Perform an image block operation on the die machine vision image to obtain an image block matrix corresponding to the die machine vision image; Perform an image embedding process on the image block matrix to obtain an image embedding matrix corresponding to the image block matrix; Determine the candidate feature representation corresponding to each image element according to the image embedding matrix of the image pixels included in each image element.
[0010] In some embodiments, any one of the X image elements is regarded as the first image element, and the distribution number of the first image element in the die machine vision image is Y, where Y is a natural number greater than 0; the determining of the candidate feature representation corresponding to each image element according to the image embedding matrix of the image pixels included in each image element includes: Obtain the feature vectors of E image pixels corresponding to the u-th presence of the first image element in the mold machine vision image, where u ≤ Y and E is a non-zero natural number; Based on a preset variable one-dimensional array and the feature vectors of the E image pixels, determine the commonality measurement coefficient corresponding to the feature vector of each image pixel among the E image pixels; Perform a classification operation on the commonality measurement coefficients corresponding to the feature vectors of each image pixel among the E image pixels to obtain the influence coefficient corresponding to the feature vector of each image pixel among the E image pixels; According to the influence coefficient corresponding to the feature vector of each image pixel among the E image pixels, perform an eccentricity adjustment process on the feature vectors of the E image pixels to obtain the u-th candidate feature representation corresponding to the first image element.
[0011] In some embodiments, any one of the X image elements is regarded as the first image element, and the first image element corresponds to Y candidate feature representations, where Y is a natural number greater than 0; determining the image element feature representation of each image element based on the adjusted candidate feature representations of each image element includes: If Y is equal to 1, then determine the adjusted candidate feature representation of the first image element as the image element feature representation of the first image element; If Y is greater than 1, then perform a merging operation on the Y adjusted candidate feature representations of the first image element to obtain the image element feature representation of the first image element.
[0012] In some embodiments, performing the merging operation on the Y adjusted candidate feature representations of the first image element to obtain the image element feature representation of the first image element includes: Perform a mean calculation operation on the Y candidate feature representations corresponding to the first image element to obtain the image element feature representation of the first image element.
[0013] In some embodiments, any two of the X image elements are regarded as the first image element and the second image element; classifying the element relationship of the X image elements in the mold machine vision image by classifying the image element feature representations of the X image elements includes: Perform a feature combination operation on the image element feature representation of the first image element and the image element feature representation of the second image element to obtain a combined feature representation; Obtain a second variable two-dimensional array and a second variable one-dimensional array, and perform a fully connected operation on the combined feature representation according to the second variable two-dimensional array and the second variable one-dimensional array to obtain a fully connected combined feature representation; Perform a classification operation on the combined feature representation after full connection to obtain confidence space information on the relationship between the first image element and the second image element; Based on the confidence space information on the relationship between the first image element and the second image element, determine the element involvement relationship between the first image element and the second image element in the die machine vision image.
[0014] In some embodiments, any two of the X image elements are regarded as the first image element and the second image element; the method further includes: According to the element involvement relationship between the first image element and the second image element in the die machine vision image, the first image element and the second image element, generate an element node relationship network corresponding to the die machine vision image; and, fuse the element node relationship network corresponding to the die machine vision image into the target image prior information library, where the element node relationship network includes the first image element, the second image element, and the element involvement relationship.
[0015] In some embodiments, the method further includes: Obtain a die machine vision image to be detected; Perform object detection on the die machine vision image to be detected to obtain the core composition content corresponding to the die machine vision image to be detected, where the core composition content includes two image elements; Search in the target image prior information library for a target element node relationship network that matches the core composition content corresponding to the die machine vision image to be detected; Generate a defect detection result for the die machine vision image to be detected according to the target element node relationship network.
[0016] On the other hand, the present application provides a computer device, including a memory and a processor, where the memory stores a computer program that can run on the processor, and the processor implements the steps in the above method when executing the program.
[0017] The beneficial effects of the present application include: The present application provides a method and device for detecting die defects based on deep learning, which obtains a machine vision image of a die. The machine vision image of the die is composed of a plurality of image regions, and the machine vision image of the die includes X image elements. The candidate feature representations of each of the X image elements are obtained. One image element corresponds to at least one candidate feature representation. According to the saliency focusing strategy, a feature fusion operation is performed on the candidate feature representations of each image element to obtain the adjusted candidate feature representation of each image element. A merging operation is performed on the adjusted candidate feature representations of each image element to obtain the image element feature representation of each image element. An element relationship classification is performed on the image element feature representations of the X image elements to determine the element involvement relationship of the X image elements in the machine vision image of the die. Based on this, performing a feature fusion operation on the candidate feature representations of each image element according to the saliency focusing strategy can optimize the candidate feature representations of each image element, so as to increase the accuracy of the obtained element involvement relationship between image elements.
[0018] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the technical solution of the present application. Brief Description of the Drawings
[0019] The accompanying drawings herein are incorporated into the specification and form a part of this specification. These drawings illustrate embodiments consistent with the present application and are used together with the specification to explain the technical solution of the present application.
[0020] Figure 1 It is a schematic flowchart of the implementation of a method for detecting die defects based on deep learning provided by an embodiment of the present application.
[0021] Figure 2 It is a schematic diagram of the hardware entity of a computer device provided by an embodiment of the present application. Detailed Embodiments
[0022] In order to make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0023] In the following description, "some embodiments" are involved, which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict. The terms "first / second / third" involved are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.
[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing this application and are not intended to limit this application.
[0025] An embodiment of the present application provides a method for detecting die defects based on deep learning, which can be executed by a processor of a computer device. Among them, the computer device can refer to devices with data processing capabilities such as servers, laptops, tablets, desktop computers, etc.
[0026] Figure 1 It is a schematic diagram of the implementation process of a method for detecting die defects based on deep learning provided by an embodiment of the present application, as Figure 1 shown, the method includes the following steps: Step S100: Obtain a die machine vision image, the die machine vision image is composed of several image regions, and the die machine vision image includes X image elements, where X>1.
[0027] In step S100, the computer device captures an image of the die through a machine vision system, that is, the die machine vision image. The die machine vision image is divided into several image regions, and each region may contain different parts or features of the die. These regions together constitute the visual representation of the entire die. The die machine vision image contains multiple (X>1) image elements. In the context of die detection, these image elements may represent different textures on the die surface. For example, some textures may be part of the normal design of the die, such as regular engraved lines or patterns; while other textures may be abnormal textures caused by manufacturing defects or wear during the use of the die. For example, assume that a die for manufacturing plastic products is being detected. In the machine vision image of this die, multiple image elements can be seen, including normal die engraved lines, possible cracks, wear areas, etc. These elements are presented in different shapes, sizes, and colors in the image and need to be further analyzed and identified through subsequent steps.
[0028] Step S200: Obtain the candidate feature representations of each of the X image elements, where each image element corresponds to no less than one candidate feature representation.
[0029] In step S200, the computer device deeply analyzes each image element obtained in step S100 to extract their features. A feature representation is an abstraction and encoding of the original data, capable of reflecting the essential attributes of the data. In the scenario of mold defect detection, features may include the thickness, direction, density, etc. of the texture, which are all important clues for determining whether the mold is normal or has defects. Specifically, the computer device uses image processing techniques and machine learning algorithms to extract features from each image element. These features can be visual information such as color, shape, texture, or more complex structural information or spatial relationships. For example, in an image of a mold surface, a specific worn area may exhibit unique texture and color, which can be used as the feature representation of this area.
[0030] Importantly, each occurrence of an image element in the image will have a corresponding candidate feature representation. This means that if an image element appears multiple times in the image, then it will have multiple candidate feature representations. These candidate feature representations can be in vector form, with each vector containing a set of feature values of the image element. For example, assume that there is a crack-shaped image element that appears twice in the mold image. Then the computer device will extract two sets of candidate feature representations for this element. Each set of feature representations may include information such as the length, width, direction, color, etc. of the crack, which are encoded into a feature vector. For instance, a feature vector may be [0.5, 0.3, 45, 120], representing the relative length, relative width, direction, and color value of the crack respectively. In this way, step S200 provides a rich data basis for subsequent feature fusion and element relationship classification, helping to accurately identify the defects of the mold.
[0031] As an implementation, in step S200, obtaining the candidate feature representations of each of the X image elements specifically includes: Step S210: Perform image block operation on the mold machine vision image to obtain the image block matrix corresponding to the mold machine vision image.
[0032] In step S210, the computer device uses a specific algorithm to divide the original mold image into several small blocks, and these small blocks are arranged according to their positions in the original image to form an image block matrix.
[0033] The purpose of the image block operation is to analyze each area of the mold surface more meticulously, so as to be able to detect possible defects more accurately. By dividing the image into small blocks, we can perform independent feature extraction and analysis on each block, which helps to discover local defects that may not be obvious in the overall image. Specifically, the computer device will first determine the size and method of block division, which usually depends on the resolution of the original image and the type of defects to be detected. For example, if fine cracks or scratches need to be detected, the size of the blocks should be relatively small to capture these subtle features.
[0034] When performing the image block operation, the computer device cuts the original image into multiple small blocks according to the preset block size. Each small block contains a part of the area in the original image, and these small blocks form a matrix according to their relative positions in the original image, that is, the image block matrix.
[0035] For example, suppose there is a high-resolution mold image. To detect possible tiny cracks on the mold surface, the image can be divided into dozens of small blocks. Each small block may only contain a small area of the mold surface, such as a scoring line or a small protrusion. Through the independent analysis of these small blocks, it can be more accurately determined which areas have potential defects. The image block operation in step S210 provides the basic data for subsequent feature extraction and defect detection. It enables the computer device to analyze each area of the mold surface more finely, thereby improving the accuracy and sensitivity of defect detection.
[0036] Step S220: Perform image embedding processing on the image block matrix to obtain the image embedding matrix corresponding to the image block matrix.
[0037] In step S220, the computer device will perform image embedding processing on the image block matrix obtained in step S210 to obtain an image embedding matrix. Image embedding processing is to transform the image data from the original pixel space to a high-dimensional feature space. This transformation can be achieved through machine learning models, especially deep learning models. In the scenario of mold defect detection, the purpose of image embedding is to extract high-level features in each image block, which can more effectively represent the image content and help with subsequent defect recognition. Specifically, the computer device uses a pre-trained deep learning model, such as a convolutional neural network (CNN), to process each image block and convert each block into a feature vector, which captures the key visual information in the block. This process is automatic, and the model will learn from a large amount of training data how to effectively extract useful features from the original pixels.
[0038] For example, assume that a deep convolutional neural network is used to process image patches. For each patch, the network outputs a feature vector of a fixed length, such as a 128-dimensional or 256-dimensional vector. This vector is a highly abstract representation of the content of the original image patch, encoding key information such as the texture, shape, and edges of the patch, which is crucial for identifying defects in the mold.
[0039] By combining the feature vectors of all image patches, we obtain an image embedding matrix. This matrix is actually a high-dimensional feature set, providing rich information for subsequent classification and recognition tasks. Generally speaking, the image embedding process in step S220 converts the original image data into a feature representation form that is easier to analyze and compare. In the scenario of mold defect detection, this conversion can significantly improve the accuracy and efficiency of defect recognition.
[0040] Step S230: Determine the candidate feature representation corresponding to each image element based on the image embedding matrix of the image pixels included in each image element.
[0041] In step S230, the computer device determines the candidate feature representations of these image elements according to the image embedding matrix of the image pixels included in each image element. Specifically, each image element consists of multiple pixels, which have been converted into an image embedding matrix in the previous steps. The image embedding matrix is actually composed of a series of feature vectors, and each feature vector corresponds to the feature representation of an image patch. Therefore, for the pixels included in a certain image element, the corresponding feature vector can be used to describe the features of this image element. To determine the candidate feature representation of the image element, the computer device aggregates these feature vectors. The aggregation method can be simple averaging, weighted averaging, max pooling, etc., depending on the actual application scenario and requirements. In this way, the computer device can extract the features related to a specific image element from the image embedding matrix to form the candidate feature representation of this image element.
[0042] For example, in the scenario of mold defect detection, assume that there is an image element representing a potential defect area on the mold surface. This area may contain multiple pixels, and each pixel has a corresponding feature vector. By aggregating these feature vectors, a feature representation representing the overall features of this potential defect area can be obtained. This feature representation can be used to determine whether there is really a defect in this area, as well as the type and severity of the defect. Generally speaking, the purpose of step S230 is to extract the features related to a specific image element from the image embedding matrix to form a candidate feature representation of this image element. This process is an important link in the mold defect detection process, which provides key feature information for subsequent classification and recognition tasks. In this way, the computer device can more accurately identify the defects on the mold surface, thereby improving production efficiency and product quality.
[0043] As an implementation manner, any one of the X image elements is regarded as the first image element, and the distribution quantity of the first image element in the mold machine vision image is Y, where Y is a natural number greater than 0; in step S230, according to the image embedding matrix of the image pixels included in each image element, the candidate feature representation corresponding to each image element is determined, which specifically includes: Step S231: Obtain the feature vectors of E image pixels corresponding to the first image element when it appears for the u-th time in the mold machine vision image, where u ≤ Y and E is a non-zero natural number.
[0044] Step S231 extracts the feature vectors of a specific image element (i.e., the first image element) from the mold machine vision image. In this step, the computer device pays attention to each occurrence of the first image element in the mold machine vision image, and for each occurrence, it will obtain the feature vectors of the image pixels included in this image element. Specifically, when the computer device processes the mold machine vision image, it first identifies the first image element (which may be a potential defect area or an area that needs special attention). Since this image element may appear multiple times in the image (assuming the number of occurrences is Y), the computer device needs to analyze each occurrence. For each occurrence of the first image element (denoted by u, where u is less than or equal to Y), the computer device determines the number of image pixels included in this image element (assuming there are E pixels) and extracts the feature vectors of these pixels.
[0045] The feature vectors are extracted from the image patches through a deep learning model (such as a convolutional neural network) in the previous steps, and they represent the high-level features of the image pixels. In the scenario of mold defect detection, these feature vectors may contain key information about pixel color, texture, shape, etc., which is crucial for subsequent defect recognition.
[0046] Taking a specific example, assume that the first image element represents a potential crack area on the die surface. This area appears three times in the die machine vision image (i.e., Y = 3). At the first appearance (u = 1), it contains 10 image pixels (i.e., E = 10). The computer device will extract the feature vectors of these 10 pixels. Each vector is a multi-dimensional array, such as [0.2, 0.5, -0.1, ..., 0.3] (this array is only for example, and the actual vector may have more dimensions and different values). These feature vectors will be used for subsequent operations such as calculating the commonality measurement coefficient, classification, and weighted summation to finally determine whether there is a real defect in this potential crack area.
[0047] Step S232: Based on the one-dimensional array of preset variables and the feature vectors of the E image pixels, determine the commonality measurement coefficient corresponding to the feature vector of each image pixel among the E image pixels.
[0048] Step S232 is a step to determine the similarity between the feature vector of the image pixel and the preset standard during the die defect detection process. The computer device uses a preset one-dimensional array of variables (or called a parameter vector) to compare with the feature vectors of the E image pixels extracted from the die machine vision image to calculate the commonality measurement coefficient between each feature vector and the preset variable. This coefficient actually reflects the similarity between the feature vector and the preset standard. Specifically, the one-dimensional array of preset variables can be regarded as a standard vector representing the surface features of a normal or expected state die. This standard vector may be obtained through machine learning algorithms based on a large amount of normal die surface image data, or manually set according to expert knowledge and experience. In the scenario of die defect detection, this standard vector usually represents the typical features of a defect-free die surface.
[0049] After the computer device obtains the feature vectors of E image pixels in the first image element (potential defect area), it will compare these feature vectors with the preset standard vectors one by one. The comparison method can be calculating the cosine similarity, Euclidean distance or other similarity measurement methods between the two vectors. Through this method, the computer device can determine the similarity degree of the feature vector of each pixel to the surface feature of the normal mold. For example, assume that the preset variable one-dimensional array is [0.1, 0.2, 0.3, 0.4], and the feature vector of a certain image pixel is [0.15, 0.18, 0.32, 0.36]. The computer device can use the cosine similarity formula to calculate the commonality measurement coefficient between these two vectors. This coefficient will represent the similarity of the pixel feature to the surface feature of the normal mold. The higher the coefficient, the more likely the pixel belongs to the normal mold surface; the lower the coefficient, the more likely the pixel belongs to the defect area. In this way, step S232 helps the computer device identify the image pixels that are inconsistent with the surface features of the normal mold, providing an important basis for further defect judgment and classification.
[0050] Step S233: Perform a classification operation on the commonality measurement coefficients corresponding to the feature vectors of each of the E image pixels to obtain the influence coefficients corresponding to the feature vectors of each of the E image pixels.
[0051] Step S233 is a step to further process the image pixel features in the mold defect detection process. In this step, the computer device will perform a classification operation on the commonality measurement coefficients calculated in step S232 to assign an influence coefficient, also known as a weight, to the feature vector of each image pixel.
[0052] The commonality measurement coefficient reflects the similarity between the feature of each image pixel and the preset standard. In step S233, the computer device classifies the image pixels according to the magnitudes of these coefficients. Generally speaking, pixels with higher commonality measurement coefficients, that is, pixels that are more similar to the preset standard, are considered more likely to belong to the normal mold surface, so their influence coefficients may be relatively low. On the contrary, pixels with lower commonality measurement coefficients, indicating that they have a greater difference from the normal features and may belong to the defect area, will be given higher influence coefficients. Specifically, in the scenario of mold defect detection, assume there is a threshold for the commonality measurement coefficient. Pixels above this threshold are considered normal, while pixels below this threshold are regarded as potential defects. The computer device will assign an influence coefficient to each pixel according to this classification. For example, normal pixels may be assigned a lower weight, such as 0.5, while potential defect pixels may be assigned a higher weight, such as 1.5.
[0053] This assignment of classification and weights helps to emphasize those pixels with greater differences from normal features in subsequent defect identification, thereby improving the accuracy of defect detection. In this way, step S233 helps the computer device to more precisely locate and identify the defect areas on the mold surface, providing strong support for subsequent defect repair or product quality control.
[0054] Step S234: According to the influence coefficients corresponding to the feature vectors of each of the E image pixels, perform an eccentric adjustment process on the feature vectors of the E image pixels to obtain the u-th candidate feature representation corresponding to the first image element.
[0055] Step S234 involves weighted processing and integration of the feature vectors of image pixels to obtain a more accurate feature representation. In this step, the computer device performs an eccentric adjustment process, that is, weighted summation, on the feature vectors of the E image pixels in the first image element according to the influence coefficients (weights) calculated in the previous steps. Specifically, in the application scenario of mold defect detection, the feature vector of each image pixel contains certain information, but due to the complexity of the mold surface and possible defects, the importance of different pixels is different. Therefore, by assigning an influence coefficient (i.e., weight) to the feature vector of each pixel, the contribution of each pixel to the feature representation can be more accurately reflected.
[0056] In step S234, the computer device weights the feature vector of each pixel according to its influence coefficient. For example, if the influence coefficient of a pixel is 1.5, then the contribution of its feature vector in the weighted summation will be greater than that of other pixels with an influence coefficient of 1. In this way, the pixel features with a higher correlation with the defect will be more prominently reflected in the final feature representation.
[0057] The specific operation of weighted summation can be to multiply the feature vector of each pixel by its corresponding influence coefficient, and then add all the weighted feature vectors to obtain a comprehensive feature representation. This comprehensive feature representation is the u-th candidate feature representation corresponding to the first image element, which more accurately reflects the features of the image element and helps with subsequent defect identification and analysis.
[0058] For example, assume that the feature vectors of three pixels are [0.1, 0.2, 0.3], [0.2, 0.3, 0.4], and [0.3, 0.4, 0.5] respectively, and their influence coefficients are 1, 1.5, and 0.5 respectively. During the weighted summation process, the feature vector of the first pixel is multiplied by 1, the feature vector of the second pixel is multiplied by 1.5, and the feature vector of the third pixel is multiplied by 0.5. Then, these three weighted feature vectors are added together, and the result is the u-th candidate feature representation of the first image element. This representation will more accurately reflect the features of the image element, providing a more powerful basis for subsequent mold defect detection.
[0059] Step S300: Perform a feature fusion operation on the candidate feature representation of each image element according to the saliency focusing strategy to obtain the adjusted candidate feature representation of each image element.
[0060] Step S300 is a step for advanced processing and fusion of the features of image elements in mold defect detection. In this step, the computer device uses the saliency focusing strategy (Attention Mechanism), also known as the attention mechanism, to perform a feature fusion operation on the candidate feature representation of each image element. The purpose of doing this is to emphasize the features that are more important for the defect detection task and suppress or ignore the irrelevant or redundant information. In the specific application scenario of mold defect detection, the attention mechanism can help the computer device notice the areas or features that are most likely to contain defects. For example, there may be some tiny cracks or depressions on the surface of the mold. These subtle features may not be prominent in the global image, but they are crucial for defect detection. Through the attention mechanism, the computer device can automatically identify and emphasize these key features. The feature fusion operation is to combine multiple candidate feature representations to form a more comprehensive and informative feature representation. During this process, the computer device will perform weighted fusion on each candidate feature representation according to the weights calculated by the attention mechanism. The feature representation with a higher weight will occupy a larger proportion in the fusion process, ensuring that important features are fully reflected in the final representation. The feature fusion operation is to optimize each candidate feature representation according to the element involvement relationship (i.e., the correlation relationship between elements) among the candidate feature representations, so that the adjusted candidate feature representation has integrity and increases the accuracy of each candidate feature representation. For example, assume that an image element has three candidate feature representations, corresponding to different feature extraction methods and scales respectively. Through the attention mechanism, the computer device calculates that the weights of these three candidate feature representations are 0.6, 0.3, and 0.1 respectively. During feature fusion, the contribution of the first candidate feature representation will be the largest, followed by the second, and finally the third. In this way, the fused feature representation will focus more on the information captured by the first candidate feature representation.
[0061] In this way, step S300 helps the computer device to more effectively utilize the feature information of image elements in mold defect detection, improving the accuracy and efficiency of detection.
[0062] As an implementation, any one of the X image elements is regarded as the first image element, the first image element corresponds to Y candidate feature representations, and any one of the Y candidate feature representations of the first image element is regarded as the u-th candidate feature representation of the first image element; both u and Y are natural numbers greater than 0, and at the same time u ≤ Y. Based on this, in step S300, a feature fusion operation is performed on the candidate feature representations of each image element according to the saliency focusing strategy to obtain the adjusted candidate feature representation of each image element, specifically including: Step S310: Obtain the commonality measurement coefficients between the u-th candidate feature representation of the first image element and each of the Y candidate feature representations of the first image element, and obtain the set of commonality measurement coefficients corresponding to the u-th candidate feature representation.
[0063] Step S310 involves evaluating the similarity or correlation between the candidate feature representations of image elements. In this step, the computer device calculates the commonality measurement coefficients between the u-th candidate feature representation of the first image element and all other candidate feature representations of this image element. The commonality measurement coefficient is a quantization index used to measure the similarity degree or correlation between two feature representations. In the application scenario of mold defect detection, this coefficient helps to identify which feature representations are similar or consistent when describing the same defect or area. Specifically, the computer device will compare the u-th candidate feature representation with the other Y - 1 candidate feature representations of the first image element one by one. This comparison process may be achieved by calculating the similarity, correlation, or other measurement methods between the two feature representations. For example, methods such as cosine similarity, Pearson correlation coefficient, or Euclidean distance can be used to measure the similarity between two feature vectors.
[0064] When calculating the commonality measurement coefficients, the computer device considers each dimension between feature representations, including information such as color, texture, shape, etc., to ensure the accuracy and comprehensiveness of the coefficients. These coefficients are then collected to form a set of commonality measurement coefficients for subsequent feature fusion operations.
[0065] For example, assume that the first image element has three candidate feature representations (Y = 3), and now we want to calculate the commonality measurement coefficients between the first candidate feature representation (u = 1) and the other two feature representations. The computer device will first extract the feature vector of the first candidate feature representation, such as [0.5, 0.3, 0.2], and then compare it with the feature vectors of the second and third candidate feature representations respectively. Through calculation, we can obtain two commonality measurement coefficients, such as 0.8 (similarity with the second feature representation) and 0.6 (similarity with the third feature representation). These two coefficients form the set of commonality measurement coefficients for the first candidate feature representation. This set will be used in subsequent steps to determine the weight of each candidate feature representation in the feature fusion process.
[0066] As an implementation manner, any one of the Y candidate feature representations of the first image element is regarded as the v-th candidate feature representation of the first image element; v is a natural number greater than 0, and at the same time v ≤ Y; In step S310, obtaining the commonality measurement coefficients between the u-th candidate feature representation of the first image element and each candidate feature representation among the Y candidate feature representations of the first image element, and obtaining the set of commonality measurement coefficients corresponding to the u-th candidate feature representation, includes: Step S311: Obtain a first variable two-dimensional array and a first variable one-dimensional array.
[0067] The first variable two-dimensional array, that is, the parameter matrix, can be regarded as a transformation matrix for converting features from one representation to another. In mold defect detection, this matrix may contain weight parameters learned from a large amount of training data, and these parameters help to extract and emphasize features related to mold defects. For example, some elements in the matrix may be trained to specifically focus on specific defects such as cracks or dents on the mold surface. The first variable one-dimensional array, that is, the parameter vector, is usually used as a bias term and plays a role in adjusting the output of neurons in a neural network. In the scenario of mold defect detection, this bias vector can help the model better adapt to different lighting conditions, shooting angles, or subtle changes on the mold surface, thereby improving the robustness of the detection.
[0068] For example, assume there is a 3x3 parameter matrix W and a parameter vector b of length 3. The matrix W may be obtained through training to extract features related to defects in the mold image, while the vector b serves as a bias term to adjust the output of the model, making it more robust to small changes in the input data.
[0069] These two parameters are obtained through machine learning algorithms, such as the training process of a deep learning network. During the training process, the model continuously adjusts these parameters to minimize the difference between the predicted result and the true result, thereby improving the accuracy of mold defect detection.
[0070] Step S312: Based on the first variable two-dimensional array, perform an integration operation on the u-th candidate feature representation and the v-th candidate feature representation of the first image element to obtain an integration operation result.
[0071] Specifically, when the computer device executes step S312, it uses the previously obtained first variable two-dimensional array (parameter matrix) to perform an integration operation on the u-th candidate feature representation and the v-th candidate feature representation of the first image element. This integration operation can be regarded as a feature fusion process, aiming to combine two or more feature representations into a more comprehensive and representative feature.
[0072] Taking mold surface defect detection as an example, assume that the u-th candidate feature representation mainly captures the texture information of the mold surface, while the v-th candidate feature representation focuses more on the edge contour information of the mold. Through the integration operation, the computer device can fuse these two aspects of information to generate a comprehensive feature representation that includes both texture and edge information.
[0073] The specific implementation method of the integration operation may vary depending on the application scenario and the selected machine learning model. In a deep learning framework, this is usually achieved through convolutional layers, fully connected layers, or specific feature fusion layers. For example, if a convolutional neural network (CNN) is used, the integration operation may be a convolution process, where the first variable two-dimensional array serves as the convolution kernel, slides on the feature map, and calculates the weighted sum to generate a new fused feature map.
[0074] In one example, assume that the u-th candidate feature representation is a tensor of shape [H, W, C1] (where H is the height, W is the width, and C1 is the number of channels), and the v-th candidate feature representation is a tensor of shape [H, W, C2]. The first variable two-dimensional array can be a convolutional kernel of shape [C_out, C1 + C2, K, K] (where C_out is the number of output channels and K is the size of the convolutional kernel). During the integration operation, first, the u-th and v-th candidate feature representations are concatenated along the channel dimension to form a tensor of shape [H, W, C1 + C2], and then the result of the integration operation of shape [H, W, C_out] is obtained by performing a convolution operation with the first variable two-dimensional array.
[0075] Step S313: Perform a non-linear transformation on the result of the integration operation, and determine a commonality metric coefficient between the u-th candidate feature representation and the v-th candidate feature representation based on the result of the integration operation after the non-linear transformation and the first variable one-dimensional array.
[0076] Step S313 involves performing a non-linear transformation on the result of the integration operation and calculating the commonality metric coefficient, which is to capture the complex relationships between features and quantify the similarity between them.
[0077] First, the computer device performs a non-linear transformation on the result of the integration operation obtained in step S312. The purpose of the non-linear transformation is to introduce more complex feature relationships, enabling the model to learn and adapt to non-linear patterns in the data. In mold defect detection, this non-linear transformation can help the model better identify minor changes or irregular shapes on the mold surface, thereby improving the sensitivity of the detection.
[0078] Specifically, the non-linear transformation can be implemented through activation functions such as ReLU (Rectified Linear Unit), Sigmoid, or Tanh. Taking the ReLU function as an example, it can set all negative values to 0 while keeping positive values unchanged. This transformation helps the model learn more sparse and effective feature representations. Next, the computer device uses the result of the integration operation after the non-linear transformation and the first variable one-dimensional array (i.e., the parameter vector) to determine the commonality metric coefficient between the u-th candidate feature representation and the v-th candidate feature representation. This coefficient reflects the similarity or correlation between the two feature representations.
[0079] In mold defect detection, the commonality metric coefficient can be regarded as the degree of synergy between two feature representations when detecting defects. If both feature representations show high responses when detecting the same type of defect, then the commonality metric coefficient between them will be large. Conversely, if the two feature representations respond to different types of defects, or one feature representation does not respond to defects, then the commonality metric coefficient between them will be small. For example, assume that there is a small crack defect on the mold surface, and the u-th candidate feature representation mainly captures the edge information of the crack, while the v-th candidate feature representation pays more attention to the internal texture of the crack. After integration operations and non-linear transformations, if both of these feature representations have strong responses at the crack defect and their response patterns are similar, then the calculated commonality metric coefficient will be high. This high coefficient will indicate that the computer device should pay more attention to the synergy of these two feature representations in subsequent detection processes.
[0080] In summary, step S313, through non-linear transformation and the calculation of the commonality metric coefficient, helps the computer device capture and quantify the similarity and correlation between different feature representations in mold defect detection more accurately. This is crucial for improving the accuracy and reliability of defect detection.
[0081] Step S314: Integrate the commonality metric coefficient between the u-th candidate feature representation and the v-th candidate feature representation into the set of commonality metric coefficients corresponding to the u-th candidate feature representation to obtain the set of commonality metric coefficients corresponding to the u-th candidate feature representation.
[0082] Step S314 integrates the calculated commonality metric coefficient into the set of commonality metric coefficients of the corresponding candidate feature representation. First, the computer device has calculated the commonality metric coefficient between the u-th candidate feature representation and the v-th candidate feature representation through the previous steps. This coefficient reflects the similarity and correlation between these two feature representations when detecting mold defects.
[0083] Next, in step S314, the computer device integrates this commonality metric coefficient into the set of commonality metric coefficients corresponding to the u-th candidate feature representation. This set may have previously included the commonality metric coefficients between the u-th candidate feature representation and other candidate feature representations. The integration operation can be a simple addition, or a more complex weighted average or screening and updating according to a certain strategy. For example, if the new commonality metric coefficient is more representative or accurate than the coefficients already in the set, then the new coefficient can replace or update the old coefficient.
[0084] Illustrated with a specific example: Assume that the \(u\)-th candidate feature mainly captures scratch defects on the mold surface, while the \(v\)-th candidate feature focuses on the rust condition of the mold. Through the previous steps, the computer device has calculated the commonality measurement coefficient between these two feature representations, such as 0.7 (the value range is between 0 and 1, indicating the similarity or correlation between the two). Now, this coefficient of 0.7 will be incorporated into the set of commonality measurement coefficients of the \(u\)-th candidate feature representation.
[0085] If the set of commonality measurement coefficients of the \(u\)-th candidate feature representation previously contained a coefficient of 0.6 with another feature representation that focuses on mold cracks, then after fusion, this set will be updated to contain two coefficients: one is the coefficient of 0.7 with the rust feature, and the other is the coefficient of 0.6 with the crack feature.
[0086] In this way, step S314 helps the computer device construct a comprehensive view of the similarity and correlation between each candidate feature representation and other feature representations. This is crucial for subsequent selection of the most representative feature combinations and improvement of the accuracy and efficiency of mold defect detection.
[0087] Step S320: Obtain preset parameters, and perform a downsampling operation on each commonality measurement coefficient in the set of commonality measurement coefficients corresponding to the \(u\)-th candidate feature representation to obtain the downsampling operation results of each commonality measurement coefficient in the set of commonality measurement coefficients corresponding to the \(u\)-th candidate feature representation.
[0088] Step S320 involves downsampling the commonality measurement coefficients to extract the most important correlation information. In this step, the computer device first obtains a preset parameter, which can be a threshold, sampling rate, or other relevant parameter, to guide the downsampling operation.
[0089] The downsampling operation, also known as sparse processing, aims to screen out the most representative coefficients from the set of commonality measurement coefficients to reduce data redundancy and noise. In the application scenario of mold defect detection, this means that the computer device will select the coefficients that are most relevant and informative for the defect detection task from the set of commonality measurement coefficients of the \(u\)-th candidate feature representation according to the preset parameter.
[0090] Specifically, if the preset parameter is a threshold, then the computer device will retain those commonality measurement coefficients that are greater than or equal to this threshold and ignore the coefficients that are less than the threshold. This ensures that only those significantly relevant feature representations will be considered in subsequent feature fusion.
[0091] For example, assume that the set of commonality measurement coefficients represented by the \(u\)-th candidate feature is \([0.8, 0.6, 0.3, 0.1]\), and the preset parameter is \(0.5\). During the downsampling operation, the computer device will retain the coefficients greater than or equal to \(0.5\), that is, \(0.8\) and \(0.6\), and ignore the coefficients less than \(0.5\), that is, \(0.3\) and \(0.1\). In this way, after the downsampling operation, the set of commonality measurement coefficients we obtain becomes \([0.8, 0.6]\), and this set will be used for subsequent standardization and weighted summation operations.
[0092] In this way, step S320 helps the computer device screen the commonality measurement coefficients before feature fusion, ensuring that only the most important correlation information is retained, thereby improving the accuracy and efficiency of feature fusion. This is particularly important in mold defect detection because it can help the device more accurately identify and locate defects on the mold.
[0093] Step S330: Perform a standardization operation on the results of the downsampling operation of each commonality measurement coefficient in the set of commonality measurement coefficients corresponding to the \(u\)-th candidate feature representation to obtain the influence coefficient cluster corresponding to the \(u\)-th candidate feature representation.
[0094] Step S330 is an important step in the process of feature fusion during mold defect detection. It involves performing a standardization operation on the downsampled commonality measurement coefficients to generate an influence coefficient cluster, also known as a weight set, for weighted summation.
[0095] In step S320, a set of important commonality measurement coefficients has been screened out through the downsampling operation. However, the numerical ranges of these coefficients may vary, and directly using them for weighted summation may cause the influence of some coefficients to be too large or too small. Therefore, the purpose of step S330 is to standardize these coefficients so that they have the same scale, thereby ensuring that each coefficient can play a reasonable role according to its importance in the subsequent weighted summation process. The standardization operation, also known as normalization, usually scales the data so that it falls into a smaller specific interval, such as \([0, 1]\) or \([-1, 1]\). In the application scenario of mold defect detection, this means that the computer device will convert the downsampled commonality measurement coefficients so that their values are all within the same range.
[0096] For example, assume that after the downsampling operation, the set of commonality measurement coefficients represented by the \(u\)-th candidate feature is \([0.8, 0.6]\). When performing the standardization operation, the computer device may use the maximum-minimum normalization method to convert these two coefficients into the range of \([0, 1]\). The specific calculation is as follows: For the coefficient \(0.8\), the normalized value is: \((0.8 - 0.6) / (0.8 - 0.6)=1\).
[0097] For the coefficient 0.6, the normalized value is: (0.6 - 0.6) / (0.8 - 0.6) = 0.
[0098] In this way, a standardized influence coefficient cluster [1, 0] is obtained. Of course, this is just a simplified example, and the actual standardization process may be more complex and require considering more factors.
[0099] Through the standardization operation in step S330, the computer device can ensure that each commonality measurement coefficient can play a reasonable role according to its importance in the subsequent weighted summation process, thereby improving the accuracy and reliability of mold defect detection.
[0100] Step S340: Perform an eccentric adjustment process on the Y candidate feature representations of the first image element according to the influence coefficient cluster corresponding to the u-th candidate feature representation, and obtain the adjusted u-th candidate feature representation of the first image element.
[0101] Step S340 uses the influence coefficient cluster (weight set) calculated in the previous steps to perform weighted summation on the candidate feature representations to achieve eccentric adjustment of the features. In the application scenario of mold defect detection, this step is crucial for accurately combining multiple feature representations to more precisely describe and identify defects on the mold. Specifically, the computer device performs weighted processing on the Y candidate feature representations of the first image element according to the influence coefficient cluster calculated in step S330. Each candidate feature representation will be assigned a weight, which reflects the importance of the feature representation in describing mold defects. The process of weighted summation is actually a linear combination of different feature representations, where the contribution of each feature representation is determined by its corresponding influence coefficient (weight).
[0102] For example, assume that the first image element has three candidate feature representations, denoted as F1, F2, and F3 respectively, and the corresponding influence coefficient cluster calculated through the previous steps is [0.5, 0.3, 0.2]. When performing the eccentric adjustment process, the computer will perform weighted summation on the feature representations according to these weights. If F1, F2, and F3 are numerical feature vectors, then the result of the weighted summation is a new feature vector, and the value of each dimension is the weighted sum of the corresponding dimension values of the original feature vectors.
[0103] Assume F1 = [1, 2, 3], F2 = [4, 5, 6], F3 = [7, 8, 9], then the new feature Fn after weighted summation is expressed as: Fn = 0.5 * F1 + 0.3 * F2 + 0.2 * F3 = 0.5 * [1, 2, 3] + 0.3 * [4, 5, 6] + 0.2 * [7, 8, 9] = [0.5×1 + 0.3×4 + 0.2×7, 0.5×2 + 0.3×5 + 0.2×8, 0.5×3 + 0.3×6 + 0.2×9] = [3.1, 4.1, 5.1].
[0104] In this way, through the eccentricity adjustment process in step S340, the computer device obtains a new feature representation that integrates the information of multiple candidate feature representations, and this new feature representation may be more accurate and comprehensive when describing die defects.
[0105] As an implementation manner, the influence coefficient cluster corresponding to the u-th candidate feature representation includes Y influence coefficients, and the Y influence coefficients correspond one-to-one to the Y candidate feature representations corresponding to the first image element; based on this, in step S340, the eccentricity adjustment process is performed on the Y candidate feature representations of the first image element according to the influence coefficient cluster corresponding to the u-th candidate feature representation to obtain the u-th candidate feature representation of the first image element after adjustment, including: Step S341: Multiply the Y candidate feature representations corresponding to the first image element by the Y influence coefficients corresponding to the u-th candidate feature representation to obtain the influence adjustment results of the Y candidate feature representations corresponding to the first image element; Step S342: Sum the influence adjustment results of the Y candidate feature representations corresponding to the first image element to obtain the u-th candidate feature representation of the first image element after adjustment.
[0106] The purpose of step S340 is to perform eccentricity adjustment on the candidate feature representations through the influence coefficient cluster to strengthen the features most relevant to the defect detection task and suppress those features that may cause interference to the detection. Steps S341 and S342 will be explained and illustrated in detail below.
[0107] Step S341 involves multiplying each candidate feature representation of the first image element by the corresponding influence coefficient to obtain a weighted result. This step is actually weighting each feature, and the magnitude of the weight reflects the importance of the feature for the final defect detection task.
[0108] Taking the crack detection on the mold surface as an example, assume that the first image element has three candidate feature representations: F1, F2, and F3, which respectively correspond to the edge information of the crack, the surface texture, and the color information. For the crack detection task, the edge information may be the most important, while the color and surface texture may be relatively less important. Therefore, through step S341, the computer device will multiply F1 by a relatively large influence coefficient, and multiply F2 and F3 by relatively small influence coefficients. In this way, the edge information feature F1 will be given a greater weight in the subsequent processing.
[0109] Specifically, if the influence coefficient of F1 is 0.6, and the influence coefficients of F2 and F3 are 0.2 respectively, then step S341 will be executed as follows: F1_weighted = F1 * 0.6, F2_weighted = F2 * 0.2, F3_weighted = F3 * 0.2. In this way, each feature representation is weighted according to its importance for the crack detection task.
[0110] Step S342 is to sum up the weighted candidate feature representations to obtain the adjusted feature representation. Continuing with the above example, the u-th adjusted candidate feature representation will be F1_weighted + F2_weighted + F3_weighted. This new feature representation contains the information of all the original features and reflects the importance of each feature for the detection task through different weights.
[0111] In this way, step S340 can help the computer device make more effective use of the feature information in the mold defect detection task, improving the accuracy and efficiency of the detection.
[0112] Step S400: Based on the adjusted candidate feature representations of each image element, determine the image element feature representation of each image element.
[0113] Step S400 determines the final feature representation of each image element based on the adjustment results of the candidate feature representations of the image elements in the previous steps. This step ensures that the feature representation of each image element can accurately reflect its key information in the mold defect detection task.
[0114] When executing step S400, the computer device first reviews the adjusted candidate feature representations of each image element. These adjusted feature representations have been weighted and reconciled according to the previous steps (such as S340) to emphasize the features related to the mold defects and suppress the irrelevant or noisy features.
[0115] Taking the detection of surface cracks in a specific mold as an example, assume that in the previous steps, the computer device has adjusted multiple candidate feature representations for each image element, such as enhancing the features related to the crack edge and weakening the features related to surface texture or lighting changes. In step S400, the computer device will combine these adjusted candidate feature representations to generate a final image element feature representation for each image element. This final feature representation may be a feature vector that contains information about all the important features in the image element, and these features have been appropriately weighted to reflect their importance in crack detection. For example, the feature vector of an image element may include eigenvalue representing the sharpness, length, direction, and contrast of the crack edge.
[0116] After determining the image element feature representation of each image element, these feature representations can be used as the input of the subsequent machine learning model to train and test the performance of the model in the mold defect detection task. In this way, step S400 provides a high-quality feature data basis for subsequent defect identification and analysis.
[0117] As an implementation, any one of the X image elements is regarded as the first image element, and the first image element corresponds to Y candidate feature representations, where Y is a natural number greater than 0; in step S400, based on the adjusted candidate feature representations of each image element, the image element feature representation of each image element is determined, which specifically includes: Step S410: If Y is equal to 1, then the adjusted candidate feature representation of the first image element is determined as the image element feature representation of the first image element; Step S420: If Y is greater than 1, then a merging operation is performed on the Y adjusted candidate feature representations of the first image element to obtain the image element feature representation of the first image element.
[0118] If an image element (taking the first image element as an example here) has only one candidate feature representation (i.e., Y is equal to 1), then directly determine the adjusted candidate feature representation as the image element feature representation of this image element. This situation may occur in some specific detection tasks. For example, when the feature information contained in the image element is relatively single, or after the previous processing steps, only one feature representation is retained.
[0119] Step S420 is for those image elements with multiple candidate feature representations (i.e., Y is greater than 1). In this case, the computer device needs to perform a merging (i.e., aggregation) operation on these adjusted candidate feature representations to obtain a unified image element feature representation. The merging operation can be simple feature concatenation, weighted average, max pooling, etc., depending on the nature of the features and the requirements of the subsequent tasks.
[0120] For example, assume that in mold defect detection, a certain image element covers a complex defect area on the mold, and this area may contain multiple features such as cracks, rust, and color changes. In the previous steps, the computer device may have generated corresponding candidate feature representations for each feature and adjusted them. In step S420, these adjusted candidate feature representations will be merged into a unified feature vector to comprehensively describe the feature information of the image element.
[0121] In this way, step S400 can ensure that each image element has a clear and comprehensive feature representation. Whether it is a single feature or a composite feature, it can be effectively captured and utilized. This provides a solid foundation for subsequent defect identification, classification, and localization.
[0122] As an implementation, in step S420, a merging operation is performed on the Y candidate feature representations after adjustment of the first image element to obtain the image element feature representation of the first image element. Specifically, it includes: performing a mean calculation operation on the Y candidate feature representations corresponding to the first image element to obtain the image element feature representation of the first image element.
[0123] In this implementation, when the first image element corresponds to multiple (Y) candidate feature representations, in order to obtain a unified image element feature representation, a mean calculation operation can be adopted.
[0124] The mean calculation operation refers to averaging each feature value of the Y candidate feature representations corresponding to the first image element. The advantage of doing this is that it can balance the influence of each feature representation and obtain a relatively robust feature representation. Specifically, if each candidate feature representation is a feature vector, then the mean calculation is to average the feature values at the same position in each vector.
[0125] For example, assume that the first image element corresponds to 3 candidate feature representations, and each feature representation is a four-dimensional vector, which are respectively represented as: Candidate feature representation 1: (F1 = [0.5, 0.8, 0.1, 0.7]).
[0126] Candidate feature representation 2: (F2 = [0.6, 0.7, 0.2, 0.8]).
[0127] Candidate feature representation 3: (F3 = [0.4, 0.9, 0.15, 0.6]).
[0128] In step S420, the computer device will perform a mean calculation on each dimension of these three feature representations. The calculation results are as follows: First - dimension mean: ((0.5 + 0.6 + 0.4) / 3 = 0.5).
[0129] Second - dimension mean: ((0.8 + 0.7 + 0.9) / 3 = 0.8).
[0130] Third - dimension mean: ((0.1 + 0.2 + 0.15) / 3 = 0.15).
[0131] Fourth - dimension mean: ((0.7 + 0.8 + 0.6) / 3 = 0.7).
[0132] Therefore, through the mean - value calculation operation, the image - element feature of the first image element is represented as: ([0.5, 0.8, 0.15, 0.7]). This feature representation synthesizes the information of the three candidate feature representations and provides a more comprehensive and robust feature basis for the subsequent detection and recognition of die defects.
[0133] Step S500: Classify the element relationships of the image - element feature representations of the X image elements to determine the element involvement relationships of the X image elements in the die machine - vision image.
[0134] In the embodiment of the present application, through step S500, the element relationships of the feature representations of the image elements are classified to determine the mutual relationships of these image elements in the die machine - vision image.
[0135] When executing step S500, the computer device analyzes the feature representation of each image element extracted in the previous steps. These feature representations may contain information about shape, texture, color, edges, etc., which are the basis for judging the relationships between image elements. The specific implementation of element - relationship classification can rely on machine - learning algorithms such as support vector machine (SVM), decision tree, random forest, or neural network. Through these algorithms, the computer device can learn and identify the potential relationship patterns between image elements.
[0136] Taking the crack detection on the mold surface as an example, assume that the mold machine vision image contains multiple image elements, and some of these image elements may represent the starting point, extension path or end point of the crack. In step S500, the computer device analyzes the feature representations of these image elements, such as the sharpness of the edge, the change of color, the continuity of texture, etc., to determine the relationship between them. Through element relationship classification, the computer device can identify which image elements are related to each other, so as to construct a network diagram of the crack distribution on the mold surface. This network diagram not only shows the overall shape of the crack, but also reveals the starting position, development direction and possible influence range of the crack. Finally, the output of step S500 may be an element relationship diagram or a set of classification labels, and this information has important guiding significance for subsequent mold defect assessment, repair plan formulation and mold quality control. In this way, step S500 provides a deeper understanding and analysis tool for mold defect detection.
[0137] As an implementation, any two of the X image elements are regarded as the first image element and the second image element; in step S500, the element relationship classification is performed on the image element feature representations of the X image elements to determine the element involvement relationship of the X image elements in the mold machine vision image, which specifically includes: Step S510: Perform a feature combination operation on the image element feature representation of the first image element and the image element feature representation of the second image element to obtain a combined feature representation.
[0138] In step S510, the computer device performs a feature combination operation on the feature representations of two image elements, that is, splices their feature vectors to obtain a new combined feature representation. Specifically, assume that there are two image elements - the first image element and the second image element, which respectively represent two different regions on the mold surface. Through the previous steps, the computer device has extracted the feature representations for these two regions, and these features may include information such as shape, texture, color, etc., and are encoded as feature vectors.
[0139] Taking a specific feature vector as an example, assume that the feature vector of the first image element is [0.5, 0.8, 0.1, 0.9], which indicates that this region has a certain specific texture and color distribution. The feature vector of the second image element may be [0.3, 0.6, 0.7, 0.2], representing the features of another region.
[0140] In step S510, the computer device concatenates these two feature vectors to form a new combined feature representation. The concatenated combined feature vector is [0.5, 0.8, 0.1, 0.9, 0.3, 0.6, 0.7, 0.2]. This new combined feature vector contains all the feature information of the two image elements, providing a data basis for subsequent determination of the relationship between these two regions.
[0141] In the embodiments of the present application, this feature combination allows for simultaneous consideration of the features of multiple regions, thereby providing a more comprehensive understanding of the condition of the mold surface. For example, if there are significant differences in color, texture, etc. between two adjacent regions, this may be a potential defect signal, such as cracks, wear, or impurities. By concatenating the feature representations of these regions, these defects can be more accurately identified and classified.
[0142] Step S520: Obtain a second variable two-dimensional array and a second variable one-dimensional array, and perform a fully connected operation on the combined feature representation according to the second variable two-dimensional array and the second variable one-dimensional array to obtain the combined feature representation after full connection.
[0143] Step S520 involves further integration and refinement of the combined feature representation. In this step, the computer device obtains two important parameters: a second variable two-dimensional array (parameter matrix) and a second variable one-dimensional array (parameter vector). These parameters are usually obtained through machine learning training and are used to perform a fully connected operation on the combined feature representation. Specifically, the fully connected operation is a process of linear transformation that multiplies the combined feature representation by the parameter matrix and the parameter vector, thereby obtaining a new, integrated combined feature representation. This process can be regarded as a weighted sum of the original features, where the values in the parameter matrix and the parameter vector are the weights.
[0144] Taking a specific example, assume that a combined feature vector [0.5, 0.8, 0.1, 0.9, 0.3, 0.6, 0.7, 0.2] is obtained in step S510. In step S520, the computer device will obtain a parameter matrix (second variable two-dimensional array), such as an 8×4 matrix, and a parameter vector (second variable one-dimensional array), such as a 4-dimensional vector. These parameters are obtained through training and are designed to capture the correlations between features and refine more meaningful features.
[0145] The fully connected operation is to multiply this 8-dimensional combined feature vector by the 8x4 parameter matrix and then add the 4-dimensional parameter vector. The result is a 4-dimensional combined feature representation after full connection, which contains the weighted combination of the original features and the adjustment of the bias term.
[0146] This process is particularly important in mold defect detection because it helps computer devices better understand and identify the complex features on the mold surface. Through the fully connected operation, we can extract the features most relevant to mold defects, thus more accurately determining whether there are defects on the mold and identifying the type and degree of the defects.
[0147] Generally speaking, the fully connected operation in step S520 is a further processing and handling of the combined feature representation. It uses the trained parameters to weight and integrate the original features to extract more meaningful feature information, providing stronger support for subsequent classification and recognition tasks.
[0148] Step S530: Perform a classification operation on the combined feature representation after full connection to obtain the confidence space information of the relationship between the first image element and the second image element.
[0149] In step S530, the computer device performs a classification operation on the combined feature representation after full connection to determine the confidence space information of the relationship between the first image element and the second image element, that is, the probability distribution of each possible relationship. Specifically, the combined feature representation after full connection has fused the feature information from the two image elements and has been further refined. This feature representation is now input into a classifier, such as a softmax classifier, for the final classification.
[0150] Taking the softmax classifier as an example, it can convert the feature representation after full connection into the probability distribution of various possible relationships. Suppose we have three possible relationships: normal area - normal area, crack starting point - crack extension, impurity area - normal area. The softmax classifier will output the probabilities of these three categories, and the sum of these probabilities is 1, representing the confidence of the computer device in the relationship between the first image element and the second image element. For example, if the probability distribution output by the softmax classifier is [0.1, 0.7, 0.2], this means that the computer device believes that the most likely relationship between these two image elements is "crack starting point - crack extension" (probability 0.7), and it is less likely to be "normal area - normal area" or "impurity area - normal area". In mold defect detection, such classification results are crucial for accurately identifying defects such as cracks and impurities on the mold. By determining the relationship between image elements, we can more precisely locate the position and type of defects, and thus timely repair or replace the mold to ensure the smooth progress of the production process.
[0151] The classification operation in step S530 uses the combined feature representation after full connection to determine the type of relationship between image elements, providing an accurate basis for subsequent defect identification and classification.
[0152] Step S540: Determine the element involvement relationship between the first image element and the second image element in the mold machine vision image based on the confidence space information regarding the relationship between the first image element and the second image element.
[0153] Based on the confidence space information obtained from the previous steps, Step S540 determines the specific relationship between the first image element and the second image element in the mold machine vision image. This step is the logical judgment link in the entire detection process, which depends on the probability distribution of various possible relationships calculated in the previous steps.
[0154] Specifically, the computer device determines the relationship between the two image elements based on the confidence space information output by the softmax classifier (or other classification algorithms), that is, the probability distribution of various relationships. For example, if the probability of a certain relationship is much higher than that of other relationships, then the computer device will determine that there is such a high-probability relationship between the two image elements.
[0155] Taking the crack detection on the mold surface as an example, assume that the first image element is the starting point of a crack and the second image element is the extended part of the crack. In Step S530, the softmax classifier has output a relatively high probability that these two elements may have a "crack starting point - crack extension" relationship. In Step S540, the computer device will determine that there is indeed a "crack starting point - crack extension" involvement relationship between these two image elements in the mold machine vision image based on this high-probability judgment.
[0156] Such a judgment is crucial for subsequent defect identification and processing. Once the relationship between the image elements is determined, the computer device can more accurately identify cracks, impurities, or other types of defects on the mold and take corresponding repair or replacement measures in a timely manner.
[0157] As an implementation, consider any two of the X image elements as the first image element and the second image element; based on this, the method further includes: generating an element node relationship network corresponding to the mold machine vision image according to the element involvement relationship between the first image element and the second image element, the first image element, and the second image element; and integrating the element node relationship network corresponding to the mold machine vision image into the target image prior information library, where the element node relationship network includes the first image element, the second image element, and the element involvement relationship.
[0158] Specifically, after the computer device has completed the analysis of the relationships between multiple image elements in the die machine vision image, it can further utilize this information to construct an element node relationship network. This relationship network is a graph structure that can clearly display the connections and involved relationships between image elements.
[0159] Specifically, the computer device regards any two of the X image elements as the first image element and the second image element, and generates a relationship network containing nodes and edges based on their involved relationships in the die machine vision image. In this relationship network, each image element is represented as a node, and the involved relationships between elements are represented as edges connecting these nodes.
[0160] Taking the crack detection on the die surface as an example, assume that the computer device has determined the relationships between a series of image elements, such as which elements are the starting points of the crack, which are the extended parts of the crack, and which are the normal areas unrelated to the crack, etc. Based on this information, the computer device can construct an element node relationship network, where the nodes represent each image element, and the edges represent the relationships between them, such as the "starting point - extension" relationship, the "crack - normal area" relationship, etc.
[0161] In addition, the computer device will also integrate this element node relationship network into a target image prior information library. This information library can be regarded as a knowledge graph of image elements, which stores a large amount of prior knowledge about image elements and their relationships. By integrating the newly generated element node relationship network into this information library, the computer device can continuously enrich and update its understanding of the die surface image elements and their relationships.
[0162] Such a fusion process is of great significance for improving the accuracy and efficiency of die defect detection. First, by constructing the element node relationship network, the computer device can more intuitively understand the relationships between image elements, thereby more accurately identifying the defects on the die. Second, by integrating the newly generated relationship network into the prior information library, the computer device can utilize historical data and prior knowledge to optimize its detection algorithm, improving the sensitivity and specificity of the detection. By constructing the element node relationship network and integrating it into the prior information library, a new analysis method based on the knowledge graph is provided for die defect detection, which helps to improve the accuracy and efficiency of the detection.
[0163] As an implementation manner, the method further includes: Step S10: Obtain the die machine vision image to be detected; Step S20: Perform target detection on the die machine vision image to be detected, and obtain the core composition content corresponding to the die machine vision image to be detected, where the core composition content includes two image elements; Step S30: Search for a target element node relationship network in the target image prior information library that matches the core composition content corresponding to the machine vision image of the mold to be detected; Step S40: Generate a defect detection result for the machine vision image of the mold to be detected based on the target element node relationship network.
[0164] In step S10, the computer device obtains a machine vision image of the mold to be detected. For example, it is completed by a camera, a scanner, or other image acquisition devices. For example, a camera on a production line can capture an image of the mold in real time and transmit it to the computer device for subsequent processing.
[0165] In step S20, the computer device can use a target detection algorithm (such as YOLO, SSD, or Faster R-CNN, etc.) to analyze the acquired image and identify the core composition content in the image. These contents usually include various parts of the mold, possible defect areas, etc. In this step, the algorithm will output two or more image elements, which are the key parts of the image and may be related to the defects of the mold.
[0166] Taking the crack on the mold as an example, the target detection algorithm may identify the starting point and the extending part of the crack as two key image elements.
[0167] In step S30, search for a target element node relationship network that matches the core composition content in the target image prior information library. The computer device searches in the previously established target image prior information library for an element node relationship network that matches the core composition content identified in step S20. This information library contains a large number of element node relationship networks of previously analyzed mold images, and each relationship network describes the relationships and involvements between image elements. By searching for a matching relationship network, the computer device can utilize prior knowledge and experience to assist in the current defect detection.
[0168] In step S40, once a matching target element node relationship network is found, the computer device can use this relationship network to generate a defect detection result. For example, if a certain image element is identified as related to a crack defect in the relationship network, the computer device will mark this element in the currently detected image and generate a corresponding defect report. This process is automated and can greatly improve the efficiency and accuracy of defect detection. In summary, through the use of historical data and knowledge in the target image prior information library in the embodiments of the present application, the computer device can more accurately identify defects on the mold, improving the sensitivity and specificity of detection. By acquiring a mold machine vision image, the mold machine vision image is composed of several image regions, and the mold machine vision image includes X image elements. The candidate feature representations of each of the X image elements are obtained. One image element corresponds to no less than one candidate feature representation. According to the saliency focusing strategy, a feature fusion operation is performed on the candidate feature representations of each image element to obtain the adjusted candidate feature representation of each image element. A merging operation is performed on the adjusted candidate feature representations of each image element to obtain the image element feature representation of each image element. An element relationship classification is performed on the image element feature representations of the X image elements to determine the element involvement relationship of the X image elements in the mold machine vision image. Based on this, performing a feature fusion operation on the candidate feature representations of each image element according to the saliency focusing strategy can optimize the candidate feature representations of each image element to increase the accuracy of the obtained element involvement relationship between image elements.
[0169] It should be noted that in the embodiments of the present application, if the above-mentioned mold defect detection method based on deep learning is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the related technology, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the embodiments of the present application. The aforementioned storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc that can store program codes. In this way, the embodiments of the present application are not limited to any specific hardware, software, or firmware, or any combination of hardware, software, and firmware.
[0170] The embodiments of the present application provide a computer device, including a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the program, it implements some or all of the steps in the above method.
[0171] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, some or all of the steps in the above method are implemented. The computer-readable storage medium can be transient or non-transient.
[0172] An embodiment of the present application provides a computer program, including computer-readable code. When the computer-readable code runs on a computer device, the processor in the computer device executes to implement some or all of the steps in the above method.
[0173] An embodiment of the present application provides a computer program product. The computer program product includes a non-transient computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, some or all of the steps in the above method are implemented. The computer program product can be specifically implemented by means of hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium. In other embodiments, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.
[0174] It should be noted here that the descriptions of the above embodiments tend to emphasize the differences between the embodiments, and their similarities or similarities can be referred to each other. The descriptions of the above device, storage medium, computer program, and computer program product embodiments are similar to the descriptions of the above method embodiments and have beneficial effects similar to those of the method embodiments. For the technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of the present application, please refer to the descriptions of the method embodiments of the present application for understanding.
[0175] Figure 2 The following is a schematic diagram of the hardware entity of a computer device provided by an embodiment of the present application. As Figure 2 shown, the hardware entity of the computer device 1000 includes: a processor 1001 and a memory 1002. Among them, the memory 1002 stores a computer program that can run on the processor 1001. When the processor 1001 executes the program, the steps in the method of any of the above embodiments are implemented.
[0176] The memory 1002 stores a computer program that can run on a processor. The memory 1002 is configured to store instructions and applications executable by the processor 1001, and can also cache data to be processed or already processed by the processor 1001 and each module in the computer device 1000 (such as image data, audio data, voice communication data, and video communication data), and can be implemented by flash memory (FLASH) or random access memory (Random Access Memory, RAM).
[0177] When the processor 1001 executes the program, it implements the steps of the deep learning-based mold defect detection method in any of the above. The processor 1001 generally controls the overall operation of the computer device 1000.
[0178] An embodiment of the present application provides a computer storage medium. The computer storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the deep learning-based mold defect detection method in any of the above embodiments.
[0179] It should be pointed out here that the descriptions of the above storage medium and device embodiments are similar to those of the above method embodiments and have similar beneficial effects to the method embodiments. For the technical details not disclosed in the storage medium and device embodiments of the present application, please refer to the descriptions of the method embodiments of the present application for understanding. The above processor can be at least one of an application specific integrated circuit (ASIC), a digital signal processor (Digital Signal Processor, DSP), a digital signal processing device (Digital Signal Processing Device, DSPD), a programmable logic device (Programmable Logic Device, PLD), a field programmable gate array (Field Programmable Gate Array, FPGA), a central processing unit (Central Processing Unit, CPU), a controller, a microcontroller, and a microprocessor. It can be understood that other electronic devices implementing the functions of the above processor are also possible, and the embodiments of the present application do not make specific limitations.
[0180] The above computer storage medium / memory may be a read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), ferromagnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM), etc.; or it may be various terminals including one or any combination of the above memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.
[0181] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, the "in one embodiment" or "in an embodiment" that appears throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics may be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the magnitude of the sequence numbers of the above steps / processes does not mean the order of execution. The order of execution of each step / process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. The serial numbers of the embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments. It should be noted that in this article, the term "comprising", "including" or any other variation thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0182] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the couplings, direct couplings, or communication connections between the various components shown or discussed can be through some interfaces. The indirect couplings or communication connections of devices or units can be electrical, mechanical, or other forms.
[0183] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units. They can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0184] In addition, each functional unit in the embodiments of the present application can be all integrated in a processing unit, or each unit can be separately used as a unit, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware, or in the form of hardware plus software functional units.
[0185] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments. The foregoing storage medium includes various media that can store program codes, such as removable storage devices, read-only memory (ROM), magnetic disks, or optical discs.
[0186] Alternatively, if the above-mentioned integrated units of the present application are implemented in the form of software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the related technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present application. The foregoing storage medium includes various media that can store program codes, such as removable storage devices, ROM, magnetic disks, or optical discs.
[0187] As described above, it is only the implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application.
Claims
1. A mold defect detection method based on deep learning, characterized in that, The method includes: Obtaining a die machine vision image, the die machine vision image being composed of a plurality of image regions, and the die machine vision image including X image elements, where X > 1; Obtaining a candidate feature representation for each of the X image elements, with each image element corresponding to at least one candidate feature representation; Performing a feature fusion operation on the candidate feature representations of each image element according to a saliency focusing strategy to obtain an adjusted candidate feature representation for each image element; Based on the adjusted candidate feature representations of each image element, determining an image element feature representation for each image element; Performing an element relationship classification on the image element feature representations of the X image elements to determine the element involvement relationship of the X image elements in the die machine vision image.
2. The method according to claim 1, wherein Regarding any one of the X image elements as a first image element, the first image element corresponding to Y candidate feature representations, and regarding any one of the Y candidate feature representations of the first image element as the u-th candidate feature representation of the first image element; both u and Y are natural numbers greater than 0, and at the same time u ≤ Y; The performing a feature fusion operation on the candidate feature representations of each image element according to a saliency focusing strategy to obtain an adjusted candidate feature representation for each image element includes: Obtaining a commonality metric coefficient between the u-th candidate feature representation of the first image element and each of the Y candidate feature representations of the first image element, and obtaining a set of commonality metric coefficients corresponding to the u-th candidate feature representation; Obtaining a preset parameter, and performing a downsampling operation on each of the commonality metric coefficients in the set of commonality metric coefficients corresponding to the u-th candidate feature representation based on the preset parameter to obtain a downsampling operation result of each of the commonality metric coefficients in the set of commonality metric coefficients corresponding to the u-th candidate feature representation; Performing a normalization operation on the downsampling operation results of each of the commonality metric coefficients in the set of commonality metric coefficients corresponding to the u-th candidate feature representation to obtain a set of influence coefficients corresponding to the u-th candidate feature representation; Performing an eccentric adjustment process on the Y candidate feature representations of the first image element according to the set of influence coefficients corresponding to the u-th candidate feature representation to obtain the adjusted u-th candidate feature representation of the first image element.
3. The method according to claim 2, characterized in that, Regarding any one of the Y candidate feature representations of the first image element as the v-th candidate feature representation of the first image element; v is a natural number greater than 0, and at the same time v ≤ Y; the obtaining a commonality metric coefficient between the u-th candidate feature representation of the first image element and each of the Y candidate feature representations of the first image element, and obtaining a set of commonality metric coefficients corresponding to the u-th candidate feature representation includes: Obtaining a first variable two-dimensional array and a first variable one-dimensional array; Performing an integration operation on the u-th candidate feature representation of the first image element and the v-th candidate feature representation of the first image element based on the first variable two-dimensional array to obtain an integration operation result; Perform a non - linear transformation on the result of the integration operation, and determine the commonality measurement coefficient between the \(u\) - th candidate feature representation and the \(v\) - th candidate feature representation based on the result of the integration operation after non - linear transformation and the one - dimensional array of the first variable; Fuse the commonality measurement coefficient between the \(u\) - th candidate feature representation and the \(v\) - th candidate feature representation into the set of commonality measurement coefficients corresponding to the \(u\) - th candidate feature representation to obtain the set of commonality measurement coefficients corresponding to the \(u\) - th candidate feature representation; The influence coefficient cluster corresponding to the \(u\) - th candidate feature representation includes \(Y\) influence coefficients, and the \(Y\) influence coefficients correspond one - to - one with the \(Y\) candidate feature representations corresponding to the first image element; The eccentric adjustment process of the \(Y\) candidate feature representations of the first image element according to the influence coefficient cluster corresponding to the \(u\) - th candidate feature representation to obtain the \(u\) - th candidate feature representation of the adjusted first image element includes: Multiply the \(Y\) candidate feature representations corresponding to the first image element by the \(Y\) influence coefficients corresponding to the \(u\) - th candidate feature representation to obtain the influence adjustment result of the \(Y\) candidate feature representations corresponding to the first image element; Sum the influence adjustment results of the \(Y\) candidate feature representations corresponding to the first image element to obtain the \(u\) - th candidate feature representation of the adjusted first image element.
4. The method according to claim 1, wherein The obtaining of the candidate feature representation of each of the \(X\) image elements includes: Perform an image block operation on the die machine vision image to obtain an image block matrix corresponding to the die machine vision image; Perform an image embedding process on the image block matrix to obtain an image embedding matrix corresponding to the image block matrix; Determine the candidate feature representation corresponding to each image element according to the image embedding matrix of the image pixels included in each image element.
5. The method according to claim 4, characterized in that, Any one of the \(X\) image elements is regarded as the first image element, and the distribution number of the first image element in the die machine vision image is \(Y\), where \(Y\) is a natural number greater than 0; The determining of the candidate feature representation corresponding to each image element according to the image embedding matrix of the image pixels included in each image element includes: Obtain the feature vectors of \(E\) image pixels corresponding to the \(u\) - th presence of the first image element in the die machine vision image, where \(u\leq Y\) and \(E\) is a non - zero natural number; Based on a preset one - dimensional array of variables and the feature vectors of the \(E\) image pixels, determine the commonality measurement coefficient corresponding to the feature vector of each of the \(E\) image pixels; Perform a classification operation on the commonality measurement coefficients corresponding to the feature vectors of each of the \(E\) image pixels to obtain the influence coefficient corresponding to the feature vector of each of the \(E\) image pixels; According to the influence coefficients corresponding to the feature vectors of each of the \(E\) image pixels, perform an eccentric adjustment process on the feature vectors of the \(E\) image pixels to obtain the \(u\) - th candidate feature representation corresponding to the first image element.
6. The method according to claim 1, wherein Any one of the X image elements is regarded as the first image element, and the first image element corresponds to Y candidate feature representations, where Y is a natural number greater than 0; Determining the image element feature representation of each image element based on the adjusted candidate feature representations of each image element includes: If Y is equal to 1, then the adjusted candidate feature representation of the first image element is determined as the image element feature representation of the first image element; If Y is greater than 1, then a merging operation is performed on the Y adjusted candidate feature representations of the first image element to obtain the image element feature representation of the first image element.
7. The method according to claim 6, wherein The merging operation is performed on the Y adjusted candidate feature representations of the first image element to obtain the image element feature representation of the first image element, including: Performing a mean calculation operation on the Y candidate feature representations corresponding to the first image element to obtain the image element feature representation of the first image element.
8. The method according to claim 1, wherein Any two of the X image elements are regarded as the first image element and the second image element; classifying the element relationships of the image element feature representations of the X image elements to determine the element involvement relationships of the X image elements in the die machine vision image includes: Performing a feature combination operation on the image element feature representation of the first image element and the image element feature representation of the second image element to obtain a combined feature representation; Obtaining a second variable two-dimensional array and a second variable one-dimensional array, and performing a fully connected operation on the combined feature representation according to the second variable two-dimensional array and the second variable one-dimensional array to obtain a fully connected combined feature representation; Performing a classification operation on the fully connected combined feature representation to obtain the confidence space information of the relationship between the first image element and the second image element; Based on the confidence space information of the relationship between the first image element and the second image element, determining the element involvement relationship between the first image element and the second image element in the die machine vision image.
9. The method according to claim 1, characterized in that, Any two of the X image elements are regarded as the first image element and the second image element; the method further includes: Generating an element node relationship network corresponding to the die machine vision image according to the element involvement relationship between the first image element and the second image element, the first image element, and the second image element; and fusing the element node relationship network corresponding to the die machine vision image into the target image prior information library, where the element node relationship network includes the first image element, the second image element, and the element involvement relationship.
10. The method according to claim 9, wherein The method further includes: Obtaining a die machine vision image to be detected; Performing object detection on the die machine vision image to be detected to obtain the core composition content corresponding to the die machine vision image to be detected, and the core composition content includes two image elements; Searching in the target image prior information library for a target element node relationship network that matches the core composition content corresponding to the die machine vision image to be detected; Generate the defect detection result of the machine vision image of the mold to be detected according to the target element node relationship network.
11. A computer device, comprising a memory and a processor, the memory storing a computer program that can run on the processor, characterized in that, When the processor executes the program, it implements the steps in the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Defect sample automatic labeling method and system based on visual attention modeling fusion
CN113256581A
Road defect detection method based on deep learning
CN118918551A
Method, system and computer programs for the automatic labeling of images for defect detection
EP4300265A1
Cited By
Defect detection method for automobile metal accessory mold
CN121280334A
A method for detecting defects in automotive metal parts molds
CN121280334B